vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [ROCm][MLA][DCP] Advertise Triton MLA non-causal multi-token DCP
- [Bugfix][Frontend] Keep is_embed when round-tripping rendered placeholders
- [Doc]: license
- Remove LICENSE inclusion from MANIFEST.in
- Qwen4Exp: QSA ring assert makes num_speculative_tokens 5..8 unreachable on all block sizes
- [Bugfix][Env] Default VLLM_DEEPEP_V2_ALLOW_HYBRID_MODE to enabled
- [Bug]: V1 spec-decode proposer never constructs the positions buffer a both-XD-RoPE drafter seeds from
- [Bug]: qwen3.8-flash-next-fp8: No available shared memory broadcast block found in 60 seconds.
- [Feature]: DeepSeek-V4-Flash-Vision-Exp (multimodal) — implementation ready, two design questions
- [Bug]: Prefix caching never hits for DeepSeek-V4-Flash on Jetson Thor (SM110) — every request cold-prefills, TTFT scales linearly with context
- Docs
- Python not yet supported