vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- fix: resolve memory profiler under-prediction for GDN/Mamba hybrid models(#52872)
- [Bugfix][DP] Ignore stale per-engine coordinator stats
- [ROCm][AMD] Kimi-K3 gfx942 / MI325X Gap and Roadmap
- [RFC]: Hybrid SSM models and last-block-replay with SpecDec and APC enabled
- [Model] Declare SupportsEagle3 on Qwen3ASRForConditionalGeneration
- [Bug]: GLM-5.2 MTP accepts 0% of drafts on MI355X (gfx950); disabling expert parallelism hits hipErrorIllegalAddress
- Add separate post-thinking sampling parameters
- Sm89
- [Bugfix][AMD] A GPU-CPU KV transfer fault in OffloadingConnector takes the engine down instead of degrading the cache.
- [Bugfix] Pad the MiniMax-M3 decode index tile to a power of two
- Docs
- Python not yet supported