vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [KV Connector] Support heterogeneous TP sharing for hybrid KV cache in Mooncake Store
- [Bug]: Cutlass int8 kernel never declines on SM120, making the Triton int8 fallback unreachable
- [Bugfix][Quantization] Decline CUTLASS int8 w8a8 on SM100+ so Triton fallback is reachable
- [Bugfix][Frontend] Treat an explicit null tool_choice as an absent one
- [Bug]: sm_120 hybrid-GDN NVFP4 dies under sustained load whenever CUDA graphs are on — persists 0.26.0 → 0.28.0, clean on 0.24.0; PIECEWISE and TRITON_ATTN both fail, only enforce_eager survives
- [RFC]: Reduced sampling for tensor-parallel decoding
- [Bugfix] Update shared expert config during elastic EP reconfiguration
- [Feature]: In the framework, there are many assert statements. How can we optimize the issue of service processes crashing due to asserts?
- [Perf][Kernel][Attention] Reorder causal-prefill launches in Triton unified attention
- [Kimi K3] Support sequence parallelism with pipeline parallelism
- Docs
- Python not yet supported