vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bug] XPU qnorm/rope kernel: int32 overflow in the address computation past 2^31 q elements
- [Quantization] Implement MXFP4 linear method
- [Bugfix][ROCm] Fix DeepSeek-V4 regression caused by moving to MRV2
- [Bugfix][Frontend] Bound Cohere stop sequences
- [MRV2][CUDA Graph] Guard attention metadata addresses before replay
- [Bug]: Harmony browser.* streaming emits both MCP and web-search events
- [Bugfix][Structured Output][Spec Decode] Validate accepted blocks before commit
- [bot] Enable Gemma4 E4B inference on Intel XPU with per-layer heterogeneous head_dim
- [bot] Enable google/gemma-4-E4B-it-assistant inference on Intel XPU
- [Bug][DP] Tie-break rotation from #47420 degrades prefix-cache locality for multi-turn traffic under internal DP LB
- Docs
- Python not yet supported