vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix][Frontend] Convert Anthropic tool_reference items to text blocks
- [Bug] Cross-node TP CUDA-graph replay deadlocks with NCCL 2.28.9 (fixed by NCCL 2.30.4) — watchdog "RPC call to sample_tokens timed out"
- [Frontend] Write run-batch responses incrementally instead of at the end
- [ci] fix wheels.vllm.ai nightly page link
- [Bug]: FlashInfer TRT-LLM bf16 MoE weight-conversion OOM recurs post-#45589 at TP2 (120B MoE, GB200) — cumem MemPool still doesn't reclaim under pressure
- [Bugfix][Frontend] Normalize reasoning_content on tokenize and pooling chat requests
- [Feature] Batch-invariant support for speculative decoding
- [Bug]: Marlin MoE output depends on equivalent within-expert token ordering
- [Metrics] Report shared-prefix tokens lost to a missing sparse-retention checkpoint
- [Bug]: Kimi-K3 CUDA graph capture silently corrupts output at batch=1; three distinct failure modes across cudagraph modes.
- Docs
- Python not yet supported