vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bug]: PLE short-conv batched prefill pads all requests to batch-max query length, causing transient activation OOM on mixed-length workloads(Qwen3.8-Flash-Next)
- [XPU] Route small TP all-reduces to a Level Zero IPC kernel
- [ROCm] Fix APU unified-memory accounting across GTT and host pressure
- [RFC] Bounded, restart-safe capacity management for the filesystem KV offload tier
- [Perf][Feat] Kimi-K3 de-JITification
- [Bug][XPU] XPU graph capture + MTP num_speculative_tokens=4 produce non-deterministic, wrong logits at temperature=0 (k<=3 clean, eager clean, bs>=2 eager fallback clean, compile-independent)
- [Bug][Core] --enable-dbo forces the V1 model runner, making dual batch overlap unreachable for every DEFAULT_V2_MODEL_RUNNER_ARCHITECTURES model (and killing startup for GLM-5.3-Flash)
- [Attention][DeepSeek-V4] Add SM90 Q8KV8 sparse MLA prefill
- [Bug]: XPU sets `use_static_cuda_launcher`, but Inductor reads `use_static_triton_launcher`
- [Bugfix][CUDA] Exclude SM121 from DeepGEMM support
- Docs
- Python not yet supported