vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix] Keep every reasoning block in the Muse Glimmer parser
- [Perf] Scale num_warps with tile size in Triton reshape-and-cache kernels
- [Bug]: MTP speculative decoding produces repetition collapse with turboquant_* KV cache on sm120 (Qwen3.8-27B hybrid GDN)
- [Bugfix][Spec Decode][Structured Output] Drive grammar masks from GPU logit counts
- [Bug]: Sleep-mode Level 2 wake does not reload speculative-decoding draft weights — silent ~2x decode slowdown, acceptance drops to 0
- [Bugfix][Spec Decode] Skip multimodal registry check for draft model configs
- [Core] Keep client stop strings dormant inside the reasoning segment
- [Bugfix] Reload speculative draft weights after Level 2 sleep wake
- [Bug]: `/v1/messages` crashes with "Unexpected item type in content" when a `tool_result` block contains a `tool_reference` item
- fix(quantization): update test_configs assertions for AWQ resolution and CPU fallback
- Docs
- Python not yet supported