vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bug]: DC unavailable for GlmMoeDsa (GLM-5.2 FP8) on SM90 — sparse MLA backend lacks decode-LSE support
- vLLM v0.26.0 + Ray TP8: intermittent wedge, `RayWorkerProc rank=[0] died unexpectedly` after `shm_broadcast` block starvation; API stays up but generation is dead
- [Bugfix][Frontend] Support Prometheus name[] metric filtering
- [7/N][KV Connector][NIXL] Stage host KV reads through device memory
- [Model] Enable LoRA support for tower and connector in GLM-ASR
- [Mamba] Support quantized FlashInfer ReplaySSM state cache
- [ROCm][Perf][MiniMax-M3] Optimize sparse GQA prefill attention
- [ROCm][MoE] Pack UE8M0 scales directly in Triton silu_mul_quant
- [ROCm][Perf] W4A16: magic-bias dequant + scale hoist for gfx90a, gated on BLOCK_M <= 16
- [ROCm][Perf] W4A16 gfx90a: extend the narrow-tile rung to M<=16
- Docs
- Python not yet supported