vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix] Register hybrid prefix-cache boundaries at hittable positions
- [Bugfix][Core] Retain align-mode mamba grid and decode-end checkpoints
- [XPU] Add env overrides for DiffKV attention 2D launch and tile size
- [Bugfix][Build] Declare fused_gdn_decode_post_conv_mtp behind the GDN guard, not KDA
- [Model] Remove native LongCat implementations
- [Frontend] Add opt-in middleware for unset chat fields
- [ROCm][Perf] Fuse MiniMax-M3 sparse cache insertion with AITER
- [ROCm][Perf][DeepSeek V4] FP8 shared-expert load-time requant for AITER fused MoE on gfx942
- [Bugfix] Fix out-of-bounds attrIdxs in cuMemcpyBatchAsync batch KV offload
- [Kernel][Perf] Fuse clamped MoE activation and UE8M0 FP8 block quantization
- Docs
- Python not yet supported