vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [CI][BugFix] Harden DeepEP NVSHMEM re-initialization in MoE layer tests
- [ROCm][Perf][Attention] Add FlyDSL gfx950 prefill attention backend
- [Bug]: _fwd_kernel_ep_scatter_2 ignores recv_x_scale_stride1 for column-major f32 scales when num_tokens >= 2
- [Bug]: UE8M0 per_token_group_quant_fp8 scales differ between the C++ and Triton fallback paths for tiny groups
- [RFC]: Sleep Mode Tensor Ownership and Recovery
- [Bugfix][MoE] Honor recv_x_scale_stride1 in ep_scatter_2 f32 scale load
- [KV Connector][NIXL] Support pipeline-parallel producers on the pull path (incl. hybrid-Mamba)
- [Bug]: tool_choice="required" not enforced with gemma4 tool parser — prose-only response returned with finish_reason="tool_calls" and empty tool_calls
- [Bugfix][Parser] Restore engine-backed parser invocation metrics
- [Bug]: torch.compile cache key does not include num_speculative_tokens, causing compiled-graph collision across different speculative-token counts
- Docs
- Python not yet supported