vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported49 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix][ROCm] Fix channelwise GPTQ MoE fallback
- [Perf][CPU] Reduce AVX2 prefill attention register spills
- [Core] Bound async GPU output synchronization
- [Bugfix] Capture an explicit off-grid max_cudagraph_capture_size
- [Bugfix][KV Connector] Use physical KV group block size in Mooncake connector
- [Hardware][PowerPC] Prioritize bfloat16 for auto dtype on PowerPC
- Revert "[Dependency] Upgrade FlashInfer version to 0.7.0" — fixes deterministic struct-output regression
- [Bug]: Triton fused MoE incorrectly indexes per-channel weight scales with per-tensor activations
- [Perf] Increase PDL overlap for Linear LoRA Shrink and Expand
- [ROCm][CI] skip the ROCm MRV1 default where MRV1 cannot serve the config
- Docs
- Python not yet supported