vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported46 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Roadmap]: PD Disaggregation with `NixlConnector` Roadmap
- [Draft] Move Harmony encoding to renderer layer and add gRPC render server
- [RFC]: KV Offloading Roadmap
- [HelionLinearBackend][1/N] Add Helion kernel for scaled_mm
- [Benchmark] Enable reproducible benchmarking with API-usage token counts
- [KV Connector][Don't merge] PoC for the async connector lookup functionality
- Fix FlexibleArgumentParser to merge JSON and dot notation arguments correctly
- [Model] Support FP8 Mamba SSM Cache
- [MLA] Separate Quant from unified_mla_attn op
- Remove all operator overrides for batch invariance
- Docs
- Python not yet supported