vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bug]: DSpark scheduling using an overly conservative input budget
- [Bug]: Voxtral-Mini-4B-Realtime permanent engine hang on a specific audio file via /v1/audio/transcriptions (empty multimodal embeddings)
- [Mamba] Add FlashInfer ReplaySSM support for MTP
- [Bugfix] Recognize CephFS as a network FS for safetensors auto-prefetch
- [Bugfix][Cosmos3-Edge] Reorder packed patches from block-major to raster before the vision tower
- [Bugfix][CPU][RISC-V] Fix cross-compilation capability detection
- [Bugfix][CPU][RISC-V] Use explicit RNE for FP16 narrowing
- [Bug]: DeepSeek-V4-Flash on RTX PRO 6000 Blackwell (SM120) emits degenerate output — identical argmax token + identical logprob at every decode position, TP and DP+EP alike, confirmed independent of environment/install history (FLASHINFER_MLA_SPARSE_DSV4)
- Support serving a model from a subfolder of its repo
- [Bugfix] Fix stale mamba/GDN states with ngram spec decode on hybrid …
- Docs
- Python not yet supported