vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported36 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- Add generic CI trace collectors
- [CI][ROCm][Disagg] Fix MiniMax-M3 WideEP DP16 idle engine sync on headless decode child nodes
- [CI][ROCm][Disagg] Fix MiniMax-M3 WideEP DP16 zero-token worker sync for all ranks (extend PR #44601)
- [CI][ROCm][Disagg] Fix WideEP DP16 MoRIIO WRITE routing with per-request multi_pod_hosts (PR #51681)
- [Bug]: mnnvl allreduce workspace init hangs 30s and leaks GPU memory on IB-only multi-node
- Revert "[Attention] Add FlashInfer XQA decode support on SM12x" (#49718)
- [Bugfix] Fix Cosmos3-Edge processor after transformers 5.15 release
- [Bugfix][Model] Fix DiffusionGemma silently freezing attention mask under CUDA graph replay
- [Attention] Allow 3D split-KV path for small-query spec-decode verify shapes
- [V1][CUDA graph] Dispatch uniform-decode batches to a padded FULL graph instead of falling to eager PIECEWISE
- Docs
- Python not yet supported