vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [XPU] Add tuned int4_w4a16 fused-MoE config for Qwen3-30B-A3B-GPTQ-Int4 on Intel Arc Pro B70
- [ROCm][Perf] Tiny-dot kernel for the Qwen MoE shared-expert gate
- [Bugfix][Spec Decode] Pass HF token to draft config
- [Performance] Reuse max query length in model runner gather
- Enable support for Terratorch models in v2 engine
- [Security] Reject reserved transfer keys in vllm_xargs
- [RFC]: Add Watchdog for Engine and Worker Monitoring
- [Bug][KV Offload][P2P] Connect and lookup waits have no deadline, leaving requests in RETRY indefinitely
- [Bug]: cu129-nightly image ships torch 2.14.0+cu130 with cu129 torchvision/torchaudio, vllm serve dies on 'operator torchvision::nms does not exist'
- Adding triton/sycl GDN toggle
- Docs
- Python not yet supported