vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [skill][ci] enhance ci skill for non-default pipeline
- [Bugfix] Ignore signals during EngineCore shutdown to avoid premature kill
- [ROCm][Perf] Use AITER tuned GEMM for the MoE router gate
- [Bugfix] Preserve FlashInfer B12x MoE runtime tensors across reload
- [Bugfix][Frontend] message-level tool declarations
- [Frontend] strict=false in response_format json_schema
- [KVConnector] Let MultiConnector compose sub-connector hits
- [Hybrid KV Cache] Retain finalized Mamba decode checkpoints for prefix caching
- [MoE][OOT] Add prepare/finalize factory registry
- [Bugfix][KV Connector][Mooncake] Avoid stale recv completion after early abort
- Docs
- Python not yet supported