vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bug]: KV Cache Offloading AssertionError after several days of uptime with DeepSeek-V4-Pro on 3×8 H100 (TP8+DP3)
- [ROCm][Perf][M3] Triton fp32 router GEMM for decode-sized M
- [Bug]: Gemma 4 MTP + NIXL PD disaggregation — draft model KV not transferred, MTP ineffective
- [Bug]: Explicit scalar CUDA graph maximum can drop a required uniform-decode shape
- [Bugfix][Compile] Preserve tensor metadata in piecewise create_concrete_args
- [Usage]: vLLM did not stop generating tokens
- [Usage] --dtype half is a performance trap on tensor-core-less Turing (GTX 16-series): fp32 is up to 9x faster for batched decode
- [DeepSeek-V4][Preview][DO NOT MERGE] Batch-invariant kernels for 0-diff RL
- [Bugfix] Chunk FP8 dummy weight initialization to bound peak memory
- [Refactor] Standardize `benchmarks/kernels/` and support multi-device tuning
- Docs
- Python not yet supported