vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Perf][KV Offload] Select the canonical load path by fragment size
- [Feature]: [CPU][GLM5Next] Add native sparse MLA / KeyPool indexer support for GLM-5.3-Flash
- Fix full KV event reporting for sparse cache hits
- [CI][Bugfix] Deflake voxtral realtime cudagraph tests
- [CI] Deflake bitsandbytes plugin tests
- [CI] Deflake HF-vs-vLLM logprobs comparisons with near-tie tolerance
- [Bugfix][Multimodal] Release consumed MRV2 CPU inputs and share local storage
- [XPU] Move weight transfer to torch.distributed where NCCL is absent
- [Bugfix][Platform] Fall back to torch when device-level NVML memory is unsupported (GB10)
- Fix: [Bug]: --fingerprint-mode=none still emits 'system_fingerprint': null in non-streaming responses
- Docs
- Python not yet supported