vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported49 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Fast Start] Add `/health` endpoint for the weight cache daemon
- [Bugfix][KV Connector][Mooncake] Apply multi-module MTP prefill backoff in P/D
- [Misc] Name each backend and its kernel block sizes in block-size errors
- [ROCm][CI] Mirror the DSv4-Flash disaggregated DP EP group on MI355
- [Warmup] Cover DSv4 sparse top-k and DFlash input-prep Triton specializations that JIT at first request
- [Bugfix][DSv4.1] Keep the compressor ring out of the null block
- [Security] Forward cache_salt into LMCache V1 keys
- [RFC]: Pluggable KV connector metrics on the Rust frontend via connector metrics descriptor
- [Metrics][Frontend] Pluggable KV connector metrics on the Rust frontend via connector metrics descriptor
- [Bug]: EngineCore worker crashes with "RuntimeError: cancelled" on first /v1/embeddings request (CPU, bge-m3)
- Docs
- Python not yet supported