vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported44 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [fix] mismatch dim during capture graph if with --gpu-memory-utilization
- [Core] Token-level DP load balancing for prefill-heavy workloads
- DeepEP LL combine_v2
- [Bugfix] Fix GDN conv + SSM state corruption with ngram spec decode
- [Feature] add mean pool (embed) feature into generate API
- [Kernel][vllm IR] Add relu2 to vLLM IR
- Enable Expert Parallel Load Balancing (EPLB) for Kimi K2.5/2.6 and Marlin Kernel
- [Test]add pytest.mark.device_type for hardcoded device strings refactoring
- [torch.compile] Add hierarchical module trace dump for FX graphs
- [Attention][TurboQuant] Optimize k8v4 decode attention with GQA head grouping
- Docs
- Python not yet supported