vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Docs]: Clarify Mamba cache mode comments in hybrid block-size alignment
- [Rust Frontend] Add LFM2 tool parser
- [ROCm][Perf] Add AITER gated QKV + RoPE + kv-cache compile fusion
- [Bugfix][CPU] Guard unavailable WNA16 operator
- [Bugfix][Kernel] Promote BF16 causal-conv operands before accumulation
- [RFC]: Adaptive prefill token budget based on scheduling pressure
- [Feature]: Support Pipeline Parallelism (PP > 1) in NixlConnector for Disaggregated KV Cache Transfer
- [Bugfix] Pad sub-block FP8 tensor-parallel shards
- [RFC] Adaptive spin grace + bounded arch waits (aarch64 WFET, x86 WAITPKG) for shm_broadcast
- [Bugfix] Align CPU offload pool size across PP/spec-decode workers
- Docs
- Python not yet supported