vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Feature]: Upgrade XGrammar to >=0.2.4 and expose max_whitespace_cnt
- [Bugfix][Reasoning] Bound-check the Hunyuan A13B side sequence
- [Bugfix][Kernel] Select batch-invariant matmul configs from runtime M
- [Bug]: Prefix caching is a no-op for nemotron_h hybrids on GPU builds; mamba-cache-mode all crashes on the CPU backend
- [CPU] Support mamba prefix-caching block tables in the CPU conv/SSM wrappers
- [Bugfix][Core] Fail cleanly for unsupported DBO backends
- [Config] Raise ValueError instead of assert for DBO all2all backend validation
- [Bug]: Cannot load an Eagle3 model, trained with Speculators
- [Feature]: Migrate MuseGlimmer reasoning/tool parsers to the Streaming Parser Engine
- [Core] Fix priority scheduling: allow waiting high-priority request to preempt running lower-priority one at max_num_seqs
- Docs
- Python not yet supported