vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Test][Tool Parser] Add streaming split-invariance test to the common suite
- [Usage]: CLS Fine-Tuned Model Fails to Launch with Multi-LoRA
- Add live scheduler knobs and in-place KV pool reinit via a dev route
- [KV Offload] Async CPU cache initialization for SimpleCPUOffloadConnector
- [Bug]: KV offloading impossible for GLM-5.3 (DSA indexer): tail_cache block_size=4 fails the hash-alignment assertion, --mamba-cache-mode align does not reach it
- [Rust][Metrics] Consume LoRA load events in the Rust frontend
- [Bugfix] FlashInfer: read the KV cache layout from attention metadata
- [Bugfix][Model] Bake real KV caches into captured fused_norm_rope under PIECEWISE (fixes silent garbage output for DSA models on multi-node TP)
- [Docs][Frontend] Add Anthropic Messages API serving guide
- [Bugfix][ROCm] Keep GLM-5.3 (GlmMoeDsaForCausalLM) on MRV2
- Docs
- Python not yet supported