vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [RFC]: Extensible GEMM Epilogue Contracts for Residual, Norm, and Gated Activation Fusion
- [RFC]: Rank-Skew-Aware Collective Attribution for Tensor-Parallel Serving
- CustomAllreduce admits cross-node TP=2 groups and dies in cudaIpcOpenMemHandle (custom_all_reduce.cuh:143 'invalid resource handle') on MNNVL GB200 NVL72
- [Bugfix][Security] Account for peak GPU usage in PyNvVideoCodec pool lease
- [Bug]: MiniMax-M3 (MXFP8) streaming: generation stops (empty delta, finish_reason=stop) right after reasoning closes, whenever a tool call should follow — non-streaming with identical payload succeeds
- [Bug]: vLLM 0.26.0 LMCache P/D receiver reports a full KV-cache hit but decoding diverges from the monolithic baseline at token 0
- [Performance]: ~18% pooling throughput regression starting with torch 2.11 (vLLM v0.20.0+)
- [Bugfix] Fix no-op num_blocks division in get_moe_wna16_block_config
- [Feature]: Optional content fallback policy for length-truncated reasoning responses
- [Docs] Clarify which deployments create a TCPStore rendezvous listener
- Docs
- Python not yet supported