vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bug]: DFlash2 draft gets 0% acceptance with --dtype float16 on XPU (bf16 works) — Qwen3.8-27B + incoai/Qwen3.8-27B-DFlash2
- [Feature][Spec Decode] Add PARD-2 parallel draft model support with dual mode
- [Feature]: Benchmark and integrate CUTLASS Lamport GEMM + AllReduce into vLLM
- [Bugfix][SM120] Scope sparse MLA warmup autotuning to attention
- [Core][Sleep Mode] Optionally release pinned host memory after wake_up
- [MoE] Optimize fused MoE token alignment and adaptive block size for fine-grained experts
- fix(openai): omit null fields in DeltaToolCall streaming response JSON
- [ROCm] Add gelu_tanh to the AITER fp8 fused MoE and zero-allocate the padded expert weights
- [Bugfix] Support GPTQ-quantized DFlash draft models in the context-KV precompute path
- [Bugfix][MoE] Warm up the Triton fused MoE GEMMs
- Docs
- Python not yet supported