vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bug]: UE8M0 per_token_group_quant_fp8 scales differ between the C++ and Triton fallback paths for tiny groups
- [RFC]: Sleep Mode Tensor Ownership and Recovery
- [Bugfix][MoE] Honor recv_x_scale_stride1 in ep_scatter_2 f32 scale load
- [KV Connector][NIXL] Support pipeline-parallel producers on the pull path (incl. hybrid-Mamba)
- [Bug]: tool_choice="required" not enforced with gemma4 tool parser — prose-only response returned with finish_reason="tool_calls" and empty tool_calls
- [Bugfix][Parser] Restore engine-backed parser invocation metrics
- [Bug]: torch.compile cache key does not include num_speculative_tokens, causing compiled-graph collision across different speculative-token counts
- [Feature]: Expose CPU vs P2P attribution for multi-tier KV restores
- [RFC]: Optimization: Take over a partial-hit block instead of CoW copy when ref_cnt == 1
- [Rust Frontend] Add OpenAI Responses API endpoint
- Docs
- Python not yet supported