vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [ROCm] TurboQuant KV cache support for Gemma 4
- [Bugfix] Recover Kimi K3 tool calls when the model omits the think close marker
- [Bugfix][Frontend] Do not report finish_reason="tool_calls" when no tool call was made
- [Bugfix][ROCm] Pass AITER MoE the rank row count, not per-expert counts
- [Performance][MRV2] Fuse vocab-parallel LM head projection for small-K prompt logprobs
- [Bugfix][FP8 linear oracle] Verify scale dtype (e8m0 vs fp32) compatibility in FP8 linear backends
- [Doc] Update extension names in incremental build guide
- [ROCm][Quantization] Match Quark exclude semantics in the shared-expert FSE check
- [Bug] DeepSeek-V4-Flash: non-deterministic output at temperature=0, rate scales with concurrency
- [Bug]: Scheduler permanently stops admitting requests once running + skipped_waiting ("deferred") reaches max_num_seqs - engine reports healthy, KV cache mostly free, only a restart recovers
- Docs
- Python not yet supported