vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Model] Qwen4Exp: fp8_e4m3 and nvfp4 KV cache on the QSA path
- [MoE] Active-parameter-aware memory profiler for MoE models
- [Kernel][Spec Decode] Add hybrid NVFP4 LM-head sampling support
- [Bugfix] Widen the QSA raw-key ring instead of asserting divisibility
- `[Performance]: Qwen3.8-Flash-Next long-prefill workload periodically starves active decode for 3-7 minutes on 2-node DGX Spark TP2`
- [Bug][SpecDecode] DFlash2 changes greedy Qwen3.8 thinking output at token 30, including K=1 and --enforce-eager
- [Attention] Portable Triton sparse-MLA fallback for SM12x (consumer Blackwell)
- [Test][RL] Add sleep/wake E2E suite with MRV1/MRV2 runner matrix
- [Bugfix][Frontend] Remove LoRA adapter from the engine on unload
- [RFC]: Support VLLM Container Snapshot with Checkpoint/Restore in Userspace (CRIU)
- Docs
- Python not yet supported