vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix][Kernel] Fix top-k selection when the radix threshold bin overflows the smem stash
- [ROCm] gfx950 top-k policy is gated on topK==1024, leaving index_topk=2048 models (GLM-5.x) on the generic 10-way split
- [Testing]Bring back test coverage that was lost due to replacing comp…
- [Bugfix][Spec Decode] Constrain a draft's attention backend to the target's KV cache layouts
- [Bug][ROCm] Startup crash (bare AssertionError) when the AITER custom-AR cutoff is below the fused allreduce+RMSNorm compile range
- [Bug][ROCm] QuickReduce converts bf16 to fp16 without saturation: activations above 65504 silently become inf
- [Bug][ROCm] gfx950 FP32 router GEMM is unreachable for DeepSeek-family models: shape (6144, 256) is listed but the weight is bf16
- [Bug]: P2P NIXL transfer exceptions leave requests pending until client timeout
- [Bugfix] Ensure deterministic tools dict serialization
- [Bugfix] Platform-aware DCP error message for ROCm
- Docs
- Python not yet supported