vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [MoE] Complete MoE Oracle Class Migration for FP8, NVFP4, MXFP8, and WNA16 (#37753)
- [Bugfix] Guard slot mapping block table loads
- [Bugfix] Reconcile the connector local hit for divergent hybrid groups
- [Bug]: Quantized embeddings fail on many model types
- GDN / Qwen3-Next decode degenerates to a single repeated token on non-Blackwell GPUs with an fp32 SSM cache
- [Bugfix] Set _enable_mm_lora on DiffusionGemma to fix image-input engine crash
- [Bug]: Qwen3.8-Flash-Next-FP8 fails to start on 4x NVIDIA A100 due to fp8e4nv unsupported in SM80
- [Bugfix][Multimodal] Fall back when an installed torchcodec fails to import
- [Bugfix] Reject a KV cache dtype that silently reinterprets the model's bits
- [Bugfix] Stop advertising fp8_e5m2 where the KV writer stores e4m3
- Docs
- Python not yet supported