vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported37 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bug]: Generation Hang after a while when using RayExecutorV2 cross two nodes with VLLM v0.24.0
- [Bugfix][Frontend] Default null Responses background and truncation
- [Bug]: Qwen3.5 DFlash speculative decoding produces repetitive/degenerate output
- [Bugfix] Bound tile-local argmax to vocab_size in samplers
- [Bugfix][DSv4] Bound token_id before the tid2eid gather in hash-MoE routing
- [Bugfix][MoE] Bound expert_map gathers on data-derived expert ids
- [BugFix] handle pre-sharded device control env var
- [ROCm] Return Kimi-K3 MLA output directly
- Fix CUDA graph profiling for long GDN warmups
- [ROCm] Skip fresh Kimi-K3 KDA state copies
- Docs
- Python not yet supported