vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix] Drop num_stages for MLA decode with fp8 KV cache on small-smem GPUs
- fix indexing for sparse MLA on AMD
- [Cleanup][MoE] Move FlashInfer MoE helpers under fused_moe
- [Bugfix] Return validation errors for malformed OpenAI request fields
- [Bugfix][Core] Skip spec padding below prefill threshold
- [Bugfix][CPU] Allow dense attention fallback for GLM DSA
- [Bug]: FP8 KV cache causes systematic decode/prefill logprob mismatch on Hopper FA3
- [Question] async scheduling defaults to on for ROCm + MTP speculative decoding, while vLLM's own ROCm CI disables that combination (#32275) — was the hang root-caused, and should the default resolution encode it?
- [Feat] merge qkv into single arg for kda_attention op
- [Bug]: benchmark_moe.py quantizes weights with current_platform.fp8_dtype() but builds the quant config with a hardcoded float8_e4m3fn
- Docs
- Python not yet supported