vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix] Set the MXFP4 persistent-kernel constraint on consumer Blackwell
- [Bug]: Align-mode prefix caching never hits (0 / 996k queries) with --scheduling-policy priority on hybrid GDN model (post-#51113)
- [Test] Refactored the routed-experts replay tests
- [Bugfix] Scan the final window in ASR audio split-point search
- [BugFix] Cache Hit Rate with KIMI + RHAI Dspark
- [Bugfix][Core] Make the attention-head / TP divisibility error actionable
- [Performance] Vectorize EPLB packing across MoE layers
- test(models): add unit test for WeightsMapper combining and pattern mapping
- [Bug]: Triton MoE and block-FP8 GEMMs mishandle the K tile — an out-of-bounds weight read, and wrong scales when a tile spans two quantization groups
- [Bug]: when setting dp=2 dsv4f0731 ,console log and metric running-request can not see both grafana and console log
- Docs
- Python not yet supported