vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix][Pooling] Ignore token stream intervals for pooling
- [Bugfix][Core] Skip cascade prefixes during deferred KV frees
- [Bugfix][DP] Make duplicate scheduler resume idempotent
- [ROCm] Fuse decode rope + MLA KV-cache write + fp8 query assembly via AITER
- [Kernel][MoE] Add tuned Triton configs for Gemma 4 on H100 and support device family fallback
- [Parser][2/2] Record completed tool calls from ParserEngine
- [RFC] Length-aware batch composition for admission scheduling — experimental evidence, fairness fix, and where it breaks
- [Bug][ROCm] GLM-5.3-Flash kpool 32-page split (hybrid block_size=1152) GPU Memory access fault on MTP decode
- [Bugfix] Fast-fail DFlash2 FP16 speculative decoding
- [Bugfix][KV Connector] Recompute after Mooncake load failures
- Docs
- Python not yet supported