vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported49 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- Rob kimi dev branch
- [RFC]: Race-free port management: pick ports at bind time, publish over an existing channel
- [Model] Extend device-side mm normalization to Qwen3VL/Qwen3.5/Qwen4Next
- [LoRA] Add fixed-weight routed LoRA reference path
- [KVConnector][MoRI-IO] Honor VLLM_MORIIO_DEFERRED_TIMEOUT_S with observability and tests
- [Bugfix] Fix structured outputs on the V2 CPU model runner
- [CI/Build][BugFix][The Rock] Apply upstream torchao fix for python 3.14 compatibility to torchao
- [Bug]: Kimi-K3 with --kv-cache-dtype fp8 is unusable on H200/Hopper — assertion demands use_prefill_query_quantization, but that flag is silently ignored on non-Blackwell devices
- [Attention] Allow num_splits > 1 on FA2 for speculative decode
- [Bug]: kernel_warmup() has no inter-rank synchronization between stages that issue real TP/EP collectives, causing startup hangs on multi-node deployments
- Docs
- Python not yet supported