vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix][LoRA] Disable encoder graphs for tower LoRA
- [Bugfix] Muse Glimmer: apply structured outputs to the to=user answer
- Fix `max_offload_tokens` assertion error.
- [Bug]: AssertionError: n_physical_experts=256 must be divisible by ep_size=3. Adjust num_redundant_experts.
- [Bugfix][Responses] Route browser streams to web search events
- [Kernel][Model] Add manual CUDA RoPE KV-cache fusion for Llama
- [Bug][ROCm] Kimi-K3 a8w4 SiTU MoE: AITER_SITUV2_A8W4 without VLLM_ROCM_USE_AITER_MOE_SITUV2_A8W4 silently produces garbage
- [Bug]: the issues about install vllm with cuda 12.6, Why hasn't anyone answered this question?
- [Bug]: AttributeError: 'ParallelLMHead' object has no attribute 'output_size_per_partition'
- [Feature][ROCm] Enable gfx1030 (RDNA2) as a recognized RDNA arch
- Docs
- Python not yet supported