vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported43 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Core] Add support for `traverse_draft_trees` for structured outputs + specdec
- [Bugfix] Fix EPLB expert mapping for DeepSeekV2 and AXK1
- Gemma4 KV-projection optimization: Remove redundant V projection on k_eq_v layers
- [Qwen3.5] Add FlashInfer GDN decode backend
- [Bugfix] Fix MoE fake output shape for TRTLLM MXFP4
- [New Model][Nvidia] Add SM12x support for DeepSeek V4 Flash with essential fixes
- [MoE Refactor] Optional router_logits argument
- [Bugfix] Flush final KV block when SimpleCPUOffload request finishes in same step as its last full block
- [Feature]: Add request-level OTel span attribute for cached prefix-cache input tokens
- [Performance]: RMSNorm op in v0.20 IR layer prevent further pytorch/triton op fusion
- Docs
- Python not yet supported