vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Perf] Parallelize TileLang mHC JIT warmup for GLM-5.3-Flash
- feat(DSv4): wait for per-layer KV loads before attention
- [XPU] Add tuned Mamba SSU configs for Intel Arc Pro B60
- [Feature]: Compact LM Head fast path for fixed-candidate-set scoring (rerankers / relevance scoring)
- [Bug]: compressed-tensors MXFP4 W4A16 is misclassified as W4A4 and dispatched to W4A4 kernel on SM100+
- [CI] Add headroom and selection coverage for H200 fast lanes
- [MOE] Emit terminal routed-expert payloads for the Rust frontend
- [Security] Handle absent tracker in LMCache MP request_finished
- [Feature][Sampling] Compact target-token-scoring fast path (V1 runner)
- [Bug]: DFlash2 draft model fails torch.compile on XPU — dynamic-shape stride assert in custom-op fake kernel (workaround: disable compile on draft class)
- Docs
- Python not yet supported