vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported37 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [CI Failure]: test_deep_ep_v2_moe FP8 output tolerance failure
- attn_res kernel latency improvements
- [Bugfix][Spec Decode] Fix misspelled EAGLE3 aux-hidden-state layer method overrides
- [Bug]: Xid 31 MMU fault (illegal write) with flashinfer_b12x MoE backend under concurrent chunked prefill (SM120, Qwen3.5-122B-A10B-NVFP4)
- [Helion] Improve fused QK config selection
- [Bugfix] Fix Marlin MoE W13 scale permutation under tensor parallelism
- [Rust Frontend] Add MiniCPM5 XML tool parser
- [Bug]: vllm-k3-toolcall repeat and repeat error
- [New Model]: Support GLM-5.2-Vision-NVFP4
- [Bugfix][KV Connector][Mooncake] Preserve HMA region strides
- Docs
- Python not yet supported