vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [RFC]: Share KV transfer planning primitives across KV connectors
- [Bugfix][ROCm] Test 1280-token GLM sparse-MLA int32 slot overflow
- [RFC]: Request level text and derender output on `/inference/v1/generate`
- [CI] Split (H200 MIG 35GB / MI300) Model Executor into four named jobs
- [CI] Split (H200 MIG 35GB / MI300) Multimodal Models (Extended Generation 1) into Audio, Vision and Omni jobs
- [XPU]Add support for weight cache on XPU device.
- [Bugfix] Abort partially-registered requests when add_request is cancelled
- [CI] Split (B200) V1 Attention into Sparse MLA, MLA, DCP, Dense/SSM, and ROCm jobs
- [CI] Split (H200 MIG 35GB / MI355 DPX) Multimodal Models (Standard) 4 into ViT CUDAGraph, Audio and Pooling + Other jobs
- [Doc]: security.md miscategorizes dev-only endpoints as production and omits /fault_tolerance/*, /metrics
- Docs
- Python not yet supported