vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [CI][torch.compile] Add E2E correctness test for RMSNorm/SiluMul + FP8-quant fusion
- [CI] Add gated smoke test for local source builds (pip install -e .)
- Dcp sparse mla fix 0.23.1
- [Feature][KV-offloading]: Host-staged RDMA for MooncakeStoreConnector requester-only ranks
- [Bugfix][CPU] Only enable ARM GELU LUT ops for models using GELU activations
- Filter sparse MLA top-k indices per DCP rank, and fix decode LSE
- feat(tpu): support single-host pipeline parallelism execution and scheduling
- fix(moe): return None in SharedExperts to support layerwise execution and AOT tracing in DeepSeek-V4
- feat(tpu): propagate worker IPs for multi-host Ray PP on TPU
- [ROCm][MLA][DCP] Advertise Triton MLA non-causal multi-token DCP
- Docs
- Python not yet supported