vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported49 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Frontend] --api-key: authenticate all routes except a liveness allowlist
- [Bug]: Hybrid Mamba/GDN P/D disaggregation is unreachable on CPU — DS conv-layout assert vs the platform's forced SD
- [Core] Count preemptions by reason
- [Frontend] Clean Client Facing 500 Report
- [Model] Support jina-ocr-v1 natively
- Realtime/transcription logprobs
- ocs: add legacy environment troubleshooting matrix for ubuntu 20.04
- [Bug] NixlConnector on MNNVL/GB200: remote engine state is released only on new-engine handshake or shutdown, deadlocking P/D prefill replacement
- [Model] Use circular buffer for DeepSeek V4 C128 state
- [Metrics] Add priority-aware labels and metrics for priority scheduling
- Docs
- Python not yet supported