vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [KV Connector] Support heterogeneous TP sharing for hybrid KV cache in Mooncake Store
- [MoE][DeepEPV2] Add PaddedStandard activation format
- [ROCm][Kimi-K3] Enable ATOM KDA sigmoid-gating and ReplaySSM kernels
- [Frontend] Propagate reasoning effort to reasoning_strength for Muse Glimmer
- [Bug]: DFlash2 + YaRN identical 1.04M prompt gets zero prefix-cache reuse while target-only reuses ~1.039M tokens
- [Model] Support Ling 3.0 D-Spark serving
- [MRV2][PP] Defer sampled-result receives
- [feat] FlashInfer CuteDSL MegaMoE integration
- [Profiler][GPU] Extend CUDA graph capture profiling to the V2 model runner
- [Spec Decode][ROCm] Add FLy: entropy-gated deferred verification for draft-model speculative decoding
- Docs
- Python not yet supported