vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix] Load checkpoint KV-cache scales in the Qwen3.5 family
- [Core] Add cache-aware admission ordering
- [MRv2] Multi-config sampler warmup for seeded native and greedy paths (#54425, #54455)
- [Bugfix] Add worker-side watchdog for wedged device kernels
- [RFC]: Preemption victim-selection extension point (least-computed-first as the first built-in)
- [Rust Frontend] Add embeddings to text and API layers
- WSL2: GPUModelRunnerV2 hard-requires UVA at init_device, so default-config GPU serving fails at startup (RuntimeError: UVA is not available)
- [RFC] Organize the Python package by architectural domain
- [RFC]: Per-KV-cache-group prefix-cache retention, with a hit-min exemption for hit-inert EAGLE sliding-window draft groups
- [Feature]: Log block pool size and per-group blocks-per-request next to "GPU KV cache size: N tokens" (N is max_concurrency x max_model_len)
- Docs
- Python not yet supported