vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported46 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [EC Connector] SHMConnector: Share Memory based EC Connector
- Added latency and throughput benchmark for beam search
- [Roadmap]: PD Disaggregation with `NixlConnector` Roadmap
- [Draft] Move Harmony encoding to renderer layer and add gRPC render server
- [RFC]: KV Offloading Roadmap
- [HelionLinearBackend][1/N] Add Helion kernel for scaled_mm
- [Benchmark] Enable reproducible benchmarking with API-usage token counts
- [KV Connector][Don't merge] PoC for the async connector lookup functionality
- Fix FlexibleArgumentParser to merge JSON and dot notation arguments correctly
- [Model] Support FP8 Mamba SSM Cache
- Docs
- Python not yet supported