vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported46 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [RFC]: Migrate bitsandbytes and GGUF quantization support to OOT plugin
- [Kernel] Add tuning script and config infrastructure for Mamba select…
- [Bugfix][Kernel] Fix int32 overflow in LoRA do_expand_kernel and do_shrink_kernel
- [Bugfix] Fix socket utilities for IPv6 dual-stack support
- Test mi300
- [Bugfix] Specify enable_kv_cache_events to prevent falsely disabling HMA in VllmConfig
- [Bugfix] Inaccurate prefix cache metrics
- [ROCm] 1st stage of enabling torch stable on ROCm.
- [Core] Use standalone autograd_cache_key for compilation dedup optimization
- [Model][Hardware][AMD][Kernel]: Enable e2e QK Norm + RoPE + KV Cache runtime fusion for Qwen3-30B-A3B on ROCM_ATTN, ROCM_AITER_FA, and ROCM_AITER_UNIFIED_ATTN
- Docs
- Python not yet supported