vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported36 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Multimodal] Add selectable audio decoding backends, including torchcodec
- fix: lazy-import model_hosting_container_standards to prevent log sup…
- [ROCm] Grant host access to VMM allocations in sleep-mode pool
- [Bugfix] Fix/gdn flashinfer fp32 qkv cast
- [Bug]: Qwen3-ASR `/v1/realtime` second utterance copies previous transcription under KV reuse
- [Core][V1] Add residual-aware SJF scheduling policy
- [Bug]: Qwen3-ASR `/v1/realtime` returns previous transcription on silence/noise under KV reuse
- [RFC] Allocate KV cache with CUDA fabric memory for cross-node MNNVL KV transfer (GB200/GB300 NVL72)
- [RFC][Reload] Replace layerwise weight transfer with modelwise transactions
- fix(deps): bump torchvision to 0.28.1 to address CVE-2026-65918
- Docs
- Python not yet supported