vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Core] Keep client stop strings dormant inside the reasoning segment
- [Bugfix] Reload speculative draft weights after Level 2 sleep wake
- [Bug]: `/v1/messages` crashes with "Unexpected item type in content" when a `tool_result` block contains a `tool_reference` item
- fix(quantization): update test_configs assertions for AWQ resolution and CPU fallback
- [Bugfix][Frontend] Convert Anthropic tool_reference items to text blocks
- [Bug] Cross-node TP CUDA-graph replay deadlocks with NCCL 2.28.9 (fixed by NCCL 2.30.4) — watchdog "RPC call to sample_tokens timed out"
- [Frontend] Write run-batch responses incrementally instead of at the end
- [ci] fix wheels.vllm.ai nightly page link
- [Bug]: FlashInfer TRT-LLM bf16 MoE weight-conversion OOM recurs post-#45589 at TP2 (120B MoE, GB200) — cumem MemPool still doesn't reclaim under pressure
- [Bugfix][Frontend] Normalize reasoning_content on tokenize and pooling chat requests
- Docs
- Python not yet supported