vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [RFC]: Lookahead-aware prefix-cache hashing for EAGLE-style draft models
- [Feature]: k3 kvcache-dtype fp8 support on hopper
- [Bugfix][Compile] Preserve independent input strides in piecewise create_concrete_args
- [Bugfix][Compile] Remove obsolete SP+PP forced +rms_norm workaround
- [ROCm] optimize memory for GLM 5.2
- [Bugfix] Use MIG handle for GPU memory lookup
- [Bugfix][KVConnector][Mooncake] Wire LookupKeyServer/Client close into store teardown
- [Feature]: Identify vLLM in the S3 client user agent, and document the S3 endpoint override
- [Feature]: Allow an explicit expert override of the startup KV capacity check
- [Bugfix][MoE] Fix padding-row handling in contiguous-layout all2all
- Docs
- Python not yet supported