vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Perf][GLM-5.2] Reuse Sparse Physical Indices via Attention Metadata with DCP
- [8/N][warmup][DSv4] Migrate MoE execution and distributed kernels
- [Bugfix] Thinking budget: stop counting at the model's natural reasoning end
- [Bugfix][ROCm][Disagg] Run DP dummy forward on every zero-token rank
- [Bugfix][V1] Mamba align: materialize a state at every boundary and drop the speculative one-block back-off
- [Bug]: FLASHINFER backend produces degenerate output for Mistral3 (Ministral-3-3B) on sm_120 with ANY kv-cache dtype; TRITON_ATTN and FLASH_ATTN correct
- [Bugfix][KVConnector] Key ExampleConnector storage on block hashes
- [Bugfix] [PD]Correct HMA enabled log condition in NIXL base_scheduler
- [RFC]: SM90 Exact Fused DSA Prefill Indexer
- [Bug][KV Offload] OffloadingConnector fs tier: multi-group MLA+DSA (DeepSeek-V4) TP=2 lookup fully misses across restart; single-group MHA hits 99.6%
- Docs
- Python not yet supported