vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Core] Add configurable sleep to shm broadcast busy loop
- [Bugfix][CPU][RISC-V] Gate oneDNN on effective RVV target
- [Bug]: MiniMax-M3 AITER sparse PA prototype corrupts output under speculative decoding with FP8 KV (ROCm)
- [Kernel][PCP] Publish replicated KV updates through PyTorch symmetric memory
- [Bug] GDN/mamba-hybrid: profiled peak activation under-predicts prefill peak; max-num-batched-tokens also sizes the CUDA-graph pool
- [ZenCPU][Test] Add zentorch/torch version pin consistency test
- [Rust Frontend] Add Anthropic Messages API request surface and count_tokens (PR 1/3)
- [Docs] Clarify Mamba cache mode comments in hybrid block-size alignment
- [Bugfix] Fix int4_w4a16 benchmark tuning on CUDA
- [RFC]: DeepSeek-V4 Performance Optimization on ROCm (Phase Two)
- Docs
- Python not yet supported