vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix] Fix structured outputs on the V2 CPU model runner
- [CI/Build][BugFix][The Rock] Apply upstream torchao fix for python 3.14 compatibility to torchao
- [Bug]: Kimi-K3 with --kv-cache-dtype fp8 is unusable on H200/Hopper — assertion demands use_prefill_query_quantization, but that flag is silently ignored on non-Blackwell devices
- [Attention] Allow num_splits > 1 on FA2 for speculative decode
- [Bug]: kernel_warmup() has no inter-rank synchronization between stages that issue real TP/EP collectives, causing startup hangs on multi-node deployments
- [CI] Fix stale Buildkite source dependencies
- [XPU] Add sequence parallelism support for DeepSeek V4
- [MM][1/n] Paged shared memory storage for mm tensor ipc.
- [sharded state] enhance validation check about output arguments
- [v1][kv-transfer] Prioritize remote KV polling in scheduler
- Docs
- Python not yet supported