vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix] Keep sub-1 req/s rates positive in bench sweep serve_workload
- [ROCm] Enable gluon paged attention decode kernel with shuffle KV cache and sinks support in AITER Flash Attention
- [Quantization] Marlin FP4: report missing CUDA custom ops in is_supported
- [Quantization] NVFP4 W4A16: fall back to weight-only emulation when Marlin is unavailable
- [Bugfix][OpenAI][Anthropic] Merge inline system messages before rendering
- [Bugfix] DeepSeek V3.2/V4: drop whitespace content deltas around too…
- [Bugfix] Fix KV cache metrics event handling
- Add FP8 W8A8 Block kernel configs for NVIDIA GeForce RTX 4090D
- [Bugfix] Disable custom all-reduce on consumer Blackwell (sm_12x)
- [Bugfix] Isolate streaming pending-state from non-streaming reasoning end check
- Docs
- Python not yet supported