vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Attention] Add B12X sparse MLA and DSA backends
- [Bugfix][V1] Register accepted spec-decode tokens before terminal cleanup
- [Feature]: Support Sentence Transformers Transformer-Pooling-Dense CrossEncoders
- [Doc] Clarify partial-block caching exception for align-mode Mamba hybrids
- [Feature][Mamba] Support batch-invariant Mamba2 prefill and recovery
- [Bug]: DP8 P2P tier startup hangs while creating secondary NIXL/UCX agents
- [Bugfix] Gate sm_100-only MLA backend tests on the capability family
- [Bug][CPU][Spec Decode]: CPU expand_kernel shim discards its output — non-greedy spec decode consumes uninitialized memory as temperature/top_k/top_p
- [Bug]: CUTLASS FP8 linear kernel is selected on Ampere while the SM80 dispatch is INT8-only
- [Bug]: nvidia/Qwen3.8-2.4T-A95B-NVFP4 serving on Ampere
- Docs
- Python not yet supported