vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Spec Decode] Support MoE DFlash2 draft models
- [V1] Fix KV cache capacity discrepancy between cold and warm starts under torch.compile (#54383)
- docs: document hybrid NIXL requirements
- [Bug]: V1 EAGLE/MTP drafter mis-addresses KV slots on M-RoPE models
- [Bugfix][Multimodal] Honor modality-scoped mm_processor_kwargs in every model that reads them
- [Bug][Quantization] Quark MXFP4 checkpoints are unloadable for multimodal models: exclude list is in checkpoint naming while vLLM matches module naming (plus 4 related name-granularity mismatches)
- [Bugfix][Model] Fix Quark FP8 block scale loading for GLM-5.3 MXFP4
- [Bugfix] Pass quant config to Phi embeddings
- [Bugfix][Reasoning] muse_glimmer: enable thinking_token_budget
- [Bugfix][Model] Cap Voxtral Realtime transcription at audio boundary
- Docs
- Python not yet supported