vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported36 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bug]: Kimi-K3 with --kv-cache-dtype fp8 is unusable on H200/Hopper — assertion demands use_prefill_query_quantization, but that flag is silently ignored on non-Blackwell devices
- [Bugfix][DSv4] Make the C128A decode topk row stride capture-stable
- [Attention] Allow num_splits > 1 on FA2 for speculative decode
- [Rust Frontend] Optimize SSE streaming hot path
- [Quantization][Humming] Support MXFP4 weight + block-FP8 activation for MoE
- [2/N] Expose HiSparse cache metrics
- [Bug]: kernel_warmup() has no inter-rank synchronization between stages that issue real TP/EP collectives, causing startup hangs on multi-node deployments
- [CI/Build] Fix strict docs build: annotate _to_serve_args parameter
- [CI] Fix stale Buildkite source dependencies
- [XPU] Add sequence parallelism support for DeepSeek V4
- Docs
- Python not yet supported