vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported42 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [kernel] Triton version of SiluAndMulWithClamp
- vLLM crashes on OpenShift (gemma4-unified-cu129)
- Refactor MoE Oracles to use base class MoEKernelOracle
- [Bugfix][Spec Decode] Derive num_speculative_tokens from a speculators draft config
- Add Thor selective state update configs
- [Feature][Frontend] Add APC prefix cache hit rate to PD usage details
- fix(logging): restore missing logs in multi API server deployments with internal DP load balancing
- [Rust Frontend] Add Cohere command reasoning parser aliases
- [Bugfix][Kernel] Clamp moe_wna16 BLOCK_SIZE_K so BLOCK_SIZE_K // group_size stays in {1,2,4,8}
- Attention backend refactor
- Docs
- Python not yet supported