vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Test] Conformance suite for KV-cache key partitioning
- [BugFix][Core] Promote parked skipped_waiting requests under saturation
- [Feature][KV cache] Support fine-grained prefix hits for sliding-window groups
- [Bugfix] Register hybrid prefix-cache boundaries at hittable positions
- [Bugfix][Core] Retain align-mode mamba grid and decode-end checkpoints
- [XPU] Add env overrides for DiffKV attention 2D launch and tile size
- [Bugfix][Build] Declare fused_gdn_decode_post_conv_mtp behind the GDN guard, not KDA
- [Model] Remove native LongCat implementations
- [Frontend] Add opt-in middleware for unset chat fields
- [ROCm][Perf] Fuse MiniMax-M3 sparse cache insertion with AITER
- Docs
- Python not yet supported