vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix] Pad FoPE sin cache with 0 so unrotated channels pass through
- [Bugfix][Determinism] Fix the batch-invariant softmax override
- [Test][Parser] Cover unpaired reasoning closers across six tool parsers
- [Bug]: Qwen3.6-27B-FP8 eventually collapses into repeated ! tokens, affecting all subsequent requests
- [XPU] add DP + EP on external lb into buildkite CI
- [Bugfix][ROCm] Saturate bf16->fp16 conversion in QuickReduce
- [Bugfix][ROCm] Keep the AITER allreduce compile range within AITER's runtime cutoff
- [Bugfix] Replace assert-as-control-flow with proper raises in KV-cache and entrypoints
- [Perf][Reasoning] Cache thinking-token history membership
- [CPU] Clamp KV-cache auto-sizing to cgroup memory limit on non-NUMA hosts too, add CI coverage
- Docs
- Python not yet supported