vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Quantization] Skip absent fused-layer shards in the GPTQ per-layer check (Gemma 4 k_eq_v)
- [Bug]: JSON/GRAMMAR structured-output schema compilation has no timeout (regex does); deeply-nested schemas cause unbounded compile time
- [Doc] Correct the Mamba rationale on needs_kv_cache_zeroing
- [Bugfix] Drop num_stages for MLA decode with fp8 KV cache on small-smem GPUs
- fix indexing for sparse MLA on AMD
- [Cleanup][MoE] Move FlashInfer MoE helpers under fused_moe
- [Bugfix] Return validation errors for malformed OpenAI request fields
- [Bugfix][Core] Skip spec padding below prefill threshold
- [Bugfix][CPU] Allow dense attention fallback for GLM DSA
- [Bug]: FP8 KV cache causes systematic decode/prefill logprob mismatch on Hopper FA3
- Docs
- Python not yet supported