vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Responses] Implement usage tracking, streaming final items, incomplete_reason, and MCP/builtin tool rendering
- [Bugfix][Spec Decode] Profile adaptive verification tail on a schedulable batch shape
- [ROCm][Kimi-K3] Fuse MXFP4 quant into KDA decode
- [ROCm][Kimi-K3] Fuse MXFP4 quant into MLA gated o_proj
- [Bug]: prefix caching + MTP still corrupts output on hybrid Mamba/GDN models in v0.28.0 (#43559 closed but unfixed)
- [Feature]: Per-request and per-session KV cache tier attribution in OpenTelemetry traces
- metrics: expose vllm:model_config_info gauge (incl. auto-configured m…
- [Bug]: [XPU] RuntimeError: causal_conv1d does not support spec-decode and non-spec tokens in the same invocation (MTP + chunked prefill)
- [RFC]: Elastic EP support in Model Runner V2
- [Bugfix][ROCm] Derive a8w4 SiTU MoE layout from AITER_SITUV2_A8W4 too
- Docs
- Python not yet supported