vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [ROCm] Calibrate gfx950 decode Top-K routing by rows × compressed-context
- [Profiler] Add per-session profiling controls to Python and Rust frontends
- [Bug]: fp8 FlashMLA sparse decode crashes at engine init with --enforce-eager on SM90 (dummy-run aten::new_empty dispatch failure, size-dependent)
- [Bug]: Async EPLB silently stops making progress — rearrangement step counter runs away negative forever, no warning/error (EP16 cross-node, use_async=True, communicator=nixl)
- [Bug]: deepep_low_latency buffer sizing ignores EPLB redundant experts — 32 GiB limit hit, documented 2048 threshold is stale
- [RFC]: Stable runtime tensor lifecycle for sleep and live weight reload
- [Experimental] Add CUDA graph discard and lazy recapture to sleep/wake
- [Spec Decode] Support Kimi-K3 MLA drafts in DFlash2
- [Bugfix][EC Connector] Fix Qwen3-Omni transfer shapes for Mooncake and CPU/NIXL
- [ROCm][DO NOT MERGE] RDNA3 (gfx1100) full inference stack — tracking branch for splitting into landable PRs
- Docs
- Python not yet supported