vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [ROCm][Perf][DeepSeek V4] FP8 shared-expert load-time requant for AITER fused MoE on gfx942
- [Bugfix] Fix out-of-bounds attrIdxs in cuMemcpyBatchAsync batch KV offload
- [Kernel][Perf] Fuse clamped MoE activation and UE8M0 FP8 block quantization
- [Bug]: TORCHINDUCTOR_COMPILE_THREADS in the environment is silently overwritten
- [MRV2] Add platform-provided runner component factory
- [ROCm][Spec Decode] Enable Kimi-K3 DSpark with pipeline parallelism / [ROCm][投机解码] 支持 Kimi-K3 DSpark 流水线并行
- [Core] Add PROMOTION_LATENCY histogram metric for tiering offload
- [Bug]: DeepSeek V3.2 encoder renders a trailing system message inside the assistant think block
- [Bugfix][Core] Log minimum cacheable prefix size
- [Bug]: test_rms_norm fails deterministically on 26-SM MIG H200 CI agents (h200-ci-2-*), passes on 16-SM agents
- Docs
- Python not yet supported