vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Core][Perf] Pass shard index bounds as module buffers in VocabParallelEmbedding
- [CI] Run the 16 root-level tests that no step collects
- [Bug]:ling-3.0-flash-fp8 gibberish
- [Bugfix] Support NoPE models on FLASHINFER_MLA_SPARSE_SM120 and validate effective topk buffer width
- [BUGFIX] Fix dflash batched token budget
- [Perf][DSV4] Use broadcast mHC pre for DeepSeek V4 DSpark
- [Bugfix][Spec Decode] DFlash2: survive spec warmup with unfilled draft buffers (embed outside compile; clamp selector gathers)
- [Attention][Spec Decode] NVFP4 KV: open the FA2 non-causal prefill wrapper for DFlash-family drafters (sm12x)
- [Bug]: ROCm spec-decode attention-metadata allowlist is enumerated by hand and has now been patched at least three times; consider letting backends declare support
- [Spec Decode][ROCm] Add FLy: entropy-gated deferred verification for draft-model speculative decoding
- Docs
- Python not yet supported