vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix] Mirror Mistral max_seq_len into max_position_embeddings
- [Bugfix] Honor batch invariance in gated normalization row tiling
- [Perf][DSv4.1] Compact large decoder prefills with a dependency-preserving halo
- [ROCm] Add head-fused HIP sparse MLA decode kernel for gfx950
- [ROCm][Perf] Interleave M4 groups for the gfx1201 M8 BF16 head
- [XPU] Bump kernels to 0.1.15.1
- [Bugfix][DSv4.1] Use 64-token sparse-MLA pages on SM120
- [CPU] Add Qwen3.8-Flash-Next CPU backend
- [XPU] limit FA kernel block size to multipleOf 64
- [Fast Start] Charge daemon-held weights against `gpu_memory_utilization`
- Docs
- Python not yet supported