vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Perf][DSV4.1] Segment FlashInfer FP8 prefill queries for faster attention tiles
- [Bugfix] Pin FlashInfer bmm_fp8 to cuBLAS on sm_12x
- [Bugfix][KV Offload] Fix partial-tail offload for hybrid models (and under DCP)
- Multimodal content-part uuid has no length bound and reaches the EngineCore block-hash pickle+SHA-256 path (uncapped sibling of the cache_salt fix)
- [Bug]: Silent garbage output on GPUs in Confidential Computing mode: V2 model runner's UVA views of pinned host memory are stale under CC (works with VLLM_USE_V2_MODEL_RUNNER=0)
- [LoRA] Add validated receiver-local adapter staging
- [Bug] gemma4 tool parser with no reasoning parser: `_preprocess_feed`'s synthetic `<|channel>` leaks into content since #47562
- [Bugfix][Frontend] Handle malformed Anthropic tool arguments
- [Bugfix][ROCm][MoE] Fix AITER MoE routing fallback to eager torch.topk for scoring_func=softmax
- [Bugfix][Parser] Preserve Gemma4 content when reasoning parsing is disabled
- Docs
- Python not yet supported