vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported35 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Draft] Investigate silent EngineCore hang in multimodal cache eviction
- Add KVCrush KV Cache compression for vLLM
- Add sparse MLA topology index policy and benchmark
- [Bug]: Qwen3-Reranker: Process Hang with `/score` Endpoint for Specific Data
- Add Shuka-1 model support (sarvamai/shuka-1)
- [Kernel][ROCm][Perf] FlyDSL decode-attention kernel for 4-bit TurboQuant KV cache
- Support DeepSeek-V4 AMD Quark NVFP4 with emulation kernel
- [Rust Frontend][RFC] Anthropic Messages API: /v1/messages and /v1/messages/count_tokens
- [ROCm][MLA][Perf] Fuse decode QK-RoPE + Q-concat + KV-concat + KV-cache write for sparse MLA
- [Performance]: MiniMax-M3 verify-as-decode re-streams the sparse KV from HBM once per draft token
- Docs
- Python not yet supported