vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- EAGLE3 + sleep-mode Level 2 + CUDA graphs: illegal memory access at draft multi-step graph replay after wake
- [Feature] Add native SM103 BF16 head-dim 512 Q1 decode routing
- [ROCm][Kimi-K3] Fuse MLA sigmoid-mul with per-token FP8 for o_proj
- [Bug]: embedding cache dtype combination finds no valid CUDA attention backend
- [ROCm] Add V1 streamed FP8 pipeline transport
- add xpu support for qwen3.8-next
- [Bugfix] Reconcile mismatched NVFP4 MoE w13 global scales
- [Bugfix] Rebuild DFlash/DSpark fused KV without empty meta placeholders
- [Bugfix][Model] Use Triton attention for Inkling on SM8x
- [Bugfix][Spec Decode] Implement SupportsPP on remaining MTP drafters (14 classes)
- Docs
- Python not yet supported