vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [ROCm][Kimi-K3] Enable ATOM KDA sigmoid-gating and ReplaySSM kernels
- [Parser] [1/2] Add vllm:tool_calls_completed_total classifier
- [Bug]: Cold AOT compilation is charged as peak activation memory, shrinking KV cache until restart
- [Bugfix] Reject incompatible mamba_ssm_cache_dtype before it crashes EngineCore
- sched: generalize has_requests() to use _push_connectors list
- [ROCm][DeepSeek V4] Enable FHMoE with DP8 over RCCL
- [Kernel] Optimize the E=60 MoE Top-K fallback on A100
- [Core][Scheduler] Batch small prefill chunks under decode load
- [Bugfix][Kernel] GDN readout: honor fp32 SSM state precision to avoid fp16 overflow → NaN
- [Bug]: GLM-5.3-Flash — ModelOpt NVFP4 checkpoints emit invalid UTF-8 byte tokens on SM120, while a compressed-tensors NVFP4 of the same model is clean
- Docs
- Python not yet supported