vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix][MRV2][LoRA] Fix prompt logprobs mapping
- [Feature][LoRA] Support per-module MM mappings for Granite Speech
- [Log][Core] Clarify shm_broadcast long-wait diagnostics
- [Rust Frontend] Add ling3 tool and reasoning parser support
- ZmqEventPublisher.publish() blocks EngineCore thread indefinitely when event queue is full, causing shm_broadcast deadlock
- [Bugfix] Harden streaming-input session lifecycle
- [Bug] SimpleCPUOffloadConnector (eager, TP=2, nightly dev1214): engine wedges at Running:0 / Waiting:3 (1 capacity + 2 deferred) — independent data point for the #45406 stranding family, with py-spy stacks from the wedged state
- [Core] KV cache: block-range eviction for live requests
- [Build] Skip phantom outputs in strict editable installs
- [Core] Bounded-memory streaming sessions: retention, in-place KV eviction, consolidation
- Docs
- Python not yet supported