vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix][Frontend] Treat an explicit null tool_choice as an absent one
- [Bug]: sm_120 hybrid-GDN NVFP4 dies under sustained load whenever CUDA graphs are on — persists 0.26.0 → 0.28.0, clean on 0.24.0; PIECEWISE and TRITON_ATTN both fail, only enforce_eager survives
- [RFC]: Reduced sampling for tensor-parallel decoding
- [Bugfix] Update shared expert config during elastic EP reconfiguration
- [Feature]: In the framework, there are many assert statements. How can we optimize the issue of service processes crashing due to asserts?
- [Perf][Kernel][Attention] Reorder causal-prefill launches in Triton unified attention
- [Kimi K3] Support sequence parallelism with pipeline parallelism
- [Bug]: [XPU] moe_wna16 AWQ fallback compares CUDA device_capability, always -1 on XPU
- Revert "[Bugfix][Spec Decode] Capture the widest uniform decode batch by default" (#50488)
- [Feature]: cannot budget KV cache per GPU when one card of a DP group is shared with another process
- Docs
- Python not yet supported