vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported49 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bug]: Mooncake connector misreads min()-collapsed cache_config.block_size as physical block size; hybrid models with mixed KV group sizes break PD transfer
- [Bugfix] Fix TorchCodec audio IO correctness
- [Bug] `ep_gather` output store overflows int32 with DeepEP v2 expanded layout (IMA in `_fwd_kernel_ep_gather`)
- [Perf][Spec Decode] Enable fused multi-step draft decode for FlashInfer trtllm-gen
- [Core][PD] Allow FULL decode graphs for one-token prompt continuations on MRv2
- Add Mooncake EP all-to-all backend
- [Kimi-K3][ROCm] Tune the recurrent KDA decode launch
- [Bugfix][Multimodal] Reject empty Qwen2-VL video samples
- [Core] Prefer a block-outermost KV cache layout when layers can be packed
- [ROCm][MLA] Cut ~21x host dispatch from the MLA metadata build
- Docs
- Python not yet supported