vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Core] Implement post-restore resume support (2/3)
- [Bug] FlashInfer CUTLASS NVFP4 MoE gives different logits for identical requests (fused finalize); `use_fused_finalize=False` is bit-stable
- [Frontend][Core] Expose container snapshot control APIs (3/3)
- [Kernel] Add VLLM_FLASHINFER_MOE_FUSED_FINALIZE to disable the nondeterministic MoE finalize
- [Bugfix][CPU][Spec Decode] Honor use_fp64_gumbel in CPU recovered-token sampling
- [Refactor] Complete the MOE oracle / linear kernel migration
- [Bugfix] Keep the multi-port DP supervisor serving through fault-tolerant recovery
- [Bugfix][Offloading] UVA offloader: release the accelerator blocks it frees
- [Bug] modelopt NVFP4 MoE: mismatched w1/w3 global scales are detected, warned about, and then used anyway
- [ROCm] Don't search MFMA-only tuning parameters on RDNA targets
- Docs
- Python not yet supported