vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported50 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Perf][DSv4.1] Stack DSpark context KV projections
- [Bugfix] Persist FlashInfer autotune cache per rank to fix the rank-0-only cache-hit deadlock
- [Frontend] Notify systemd (READY=1) once the HTTP server is serving
- [Bugfix] Invalidate KDA merged conv-weight cache on weight reload
- [Bugfix] Skip FP8-dequant fast paths for PP-missing layers in GLM-5.3-Flash load
- [ROCm][Kimi-K3][Perf] Fuse MLA decode KV-cache write and Q-prep via AITER
- [Bugfix][ROCm] Keep gfx942 bf16 persistent MLA decode at qlen <= 4
- [KV Offload] Expose the processed prefix at request finalization
- [RFC]: Declarative cross-platform capability negotiation and effective configuration reporting
- [KV-Offloading][TP] : Expand replicated_layout detection to multi-group MLA
- Docs
- Python not yet supported