vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported49 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Test] Cover cumulative parallel sampling outputs
- [Bugfix][ROCm] Gate MXFP8 linear LDS budget on gfx950, not the Triton version string
- [CI Failure]: (DGX) Spark GPQA Eval (GPT-OSS) test_gpqa_corectness
- [Bug][CPU][Spec Decode]: PR #56323 left 6 CPU Triton fallbacks patching module globals that are no longer read — spec decode broken on arm64
- [Bugfix][ROCm][Determinism] mean_kernel: 2-D grid so huge program counts don't kill the HIP launch
- docs: explain Qwen3 parser boundary tokens in custom grammars
- [Perf] DiffusionGemma: one-pass sampler statistics kernel
- [Bug]: an unclosed <tool_call> opener written in prose holds the Qwen3 parser in tool state and swallows a later genuine tool call
- [Bugfix][MiMo] Fix ViT attention-sink semantics and omni wrapper packed_modules_mapping
- [Bug]: An error preventing the compilation of certain code about CUDA kernels soccurred with FlashInfer version 0.6.18, which was automatically installed alongside the precompiled vLLM version 0.29.0
- Docs
- Python not yet supported