vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bug]: TORCHINDUCTOR_COMPILE_THREADS in the environment is silently overwritten
- [MRV2] Add platform-provided runner component factory
- [ROCm][Spec Decode] Enable Kimi-K3 DSpark with pipeline parallelism / [ROCm][投机解码] 支持 Kimi-K3 DSpark 流水线并行
- [Core] Add PROMOTION_LATENCY histogram metric for tiering offload
- [Bug]: DeepSeek V3.2 encoder renders a trailing system message inside the assistant think block
- [Bugfix][Core] Log minimum cacheable prefix size
- [Bug]: test_rms_norm fails deterministically on 26-SM MIG H200 CI agents (h200-ci-2-*), passes on 16-SM agents
- [Bug]: XPU TP>1 consumes host RAM equal to total VRAM; workers are never isolated to their own device
- [ROCm] Tests confirmed skipped due to num_devices/num_gpus mismatch (nightly-log-verified)
- [Bug]: Out-of-bounds attrIdxs in the C++ batch memcpy path (cache_kernels.cu)
- Docs
- Python not yet supported