vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported50 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix][HiSparse] Account for GPU copies during prefix admission
- [Perf][Attention] DSA candidate-block kernels walk each row's live context: 23.63 s to 22.52 s per 1000 requests on 8 x B200 at TP=8
- [ROCm][DSv4.1] Support the NVFP4 compressed KV cache on gfx950
- [Kimi] Fuse MLP add into next attn-res kernel
- [XPU] use fused DFlash2 grouped convolution
- [Bugfix][Offload] Respect --cpu-offload-gb budget for UVA parameters
- [Perf][KV Offload] Publish async lookup results per request group
- [ROCm][DSv4.1][Perf] Fused router gate for gfx950
- [RFC]: Bound frontend drain latency for bulk aborts and long non-streaming completions
- [ROCm] Optional non-expandable CUDA graph pool to skip the AITER all-reduce copy-in
- Docs
- Python not yet supported