vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported49 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bug]: Async EPLB silently stops making progress — rearrangement step counter runs away negative forever, no warning/error (EP16 cross-node, use_async=True, communicator=nixl)
- [Bug]: deepep_low_latency buffer sizing ignores EPLB redundant experts — 32 GiB limit hit, documented 2048 threshold is stale
- [RFC]: Stable runtime tensor lifecycle for sleep and live weight reload
- [Experimental] Add CUDA graph discard and lazy recapture to sleep/wake
- [Spec Decode] Support Kimi-K3 MLA drafts in DFlash2
- [Bugfix][EC Connector] Fix Qwen3-Omni transfer shapes for Mooncake and CPU/NIXL
- [ROCm][DO NOT MERGE] RDNA3 (gfx1100) full inference stack — tracking branch for splitting into landable PRs
- [MM] Enable device normalization for Llama Nemotron VL Embed/Rerank
- [Bug]: --kv-cache-memory suggestion double-counts CUDAGraph memory (regression of #37426, reintroduced by #49208)
- [Bug]: FunctionGemma rejects hyphenated tool names only in non-streaming responses
- Docs
- Python not yet supported