vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Perf][Frontend] Render plain completion stream deltas from a template
- [Bug]: Recurring EngineCore fatal errors on DGX Spark (GB10 / sm_121): CUDA kernel launch failures after 1.5–3 days of uptime (Triton JIT + FlashInfer cuDNN FP8 GEMM)
- [Bugfix][Perf] Let speculative decode use the 3D split-KV attention path
- [Perf][LoRA] Hoist the fused MoE-LoRA shrink reduction out of the K loop
- [Bugfix][Spec Decode] DFlash2: accept unquantized linear LM heads in the candidate selector
- [Rust Frontend] Recover incomplete Kimi K3 calls at outer boundaries
- [Bug]: causal_conv1d Triton kernels round BF16 products before FP32 accumulation
- [ROCm][Perf] Kimi-K3 Mixed-dtype Biased Grouped Top-k
- [Bugfix][Hardware][AMD] Skip MiniMax-M3 AITER sparse PA under spec decode
- [Bug]: deepep_v2 crash in cudagraph mode
- Docs
- Python not yet supported