vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [PD][HeteroArch] Enable nixl kv transfer with heterogeneous buffer device
- [Kimi K3][Kernel] Fuse BF16 shared experts into latent MegaMoE tail
- [Core][Feat] Pluggable KVCacheConfigBuilder for platform/model-specific KV cache planning
- [Perf][GLM-5.2] Reuse Sparse Physical Indices via Attention Metadata with DCP
- [8/N][warmup][DSv4] Migrate MoE execution and distributed kernels
- [Bugfix] Thinking budget: stop counting at the model's natural reasoning end
- [Bugfix][ROCm][Disagg] Run DP dummy forward on every zero-token rank
- [Bugfix][V1] Mamba align: materialize a state at every boundary and drop the speculative one-block back-off
- [Bug]: FLASHINFER backend produces degenerate output for Mistral3 (Ministral-3-3B) on sm_120 with ANY kv-cache dtype; TRITON_ATTN and FLASH_ATTN correct
- [Bugfix][KVConnector] Key ExampleConnector storage on block hashes
- Docs
- Python not yet supported