vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- Add NIXL side channel advertize address for NAT'd deployments
- [Bug]: Prefill misdispatched into spec-decode FULL cudagraph when prompt length == 1 + num_speculative_tokens → silent GDN state loss, garbage output (hybrid/Qwen3-Next models)
- [Bugfix][V1] Don't crash the engine on an aborted request
- [Bugfix][KV Offload] Bound primary HIT_PENDING waits
- fix(deepseek_v4): gate head_dim==512 compressor path on has_cutedsl()
- fix(security): use segment-boundary-aware root_path stripping in auth middleware
- [Quant] Add Humming backend support to CT WNA16 MoE
- [Bugfix][ROCm] Match gemm_with_dynamic_quant_fake argument order to the real op
- [Core] Make the shm message-queue busy-wait window configurable
- Scheduler.schedule() is the most complex function in the codebase (complexity 136, 841 lines)
- Docs
- Python not yet supported