vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Kernel][Model] Add manual CUDA RoPE KV-cache fusion for Llama
- [Bug]: Client stop strings match inside the reasoning segment for think-in-prompt models, truncating CoT and yielding content=None
- [Bug][ROCm] Kimi-K3 a8w4 SiTU MoE: AITER_SITUV2_A8W4 without VLLM_ROCM_USE_AITER_MOE_SITUV2_A8W4 silently produces garbage
- [Bugfix] Raise ValueError instead of bare assert for expert-parallel divisibility check
- [Bug]: the issues about install vllm with cuda 12.6, Why hasn't anyone answered this question?
- [Bug]: AttributeError: 'ParallelLMHead' object has no attribute 'output_size_per_partition'
- [Feature][ROCm] Enable gfx1030 (RDNA2) as a recognized RDNA arch
- [Bugfix] Enforce deterministic cuDNN convolution under VLLM_BATCH_INVARIANT
- [Core][Perf] Zero new KV blocks by cache group
- # [kimik3][ROCm] Enable torch.compile for so post-grad fusion passes work (aiter::fused_qk_rmsnorm_kernel, aiter::allreduce_fusion_kernel_1stage)
- Docs
- Python not yet supported