vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Perf][Qwen] Overlap GDN QKVZ and BA input projections
- [Bugfix] Detect multi-branch allOf for xgrammar structured-output fallback
- [Perf][DSv4.1] Skip candidate sorting when all blocks fit
- [Bugfix] Defer xgrammar annotations to fix startup crash when xgrammar is unavailable
- [Performance]: GLM-5.3-Flash on H100: auto-selected FLASHINFER_MLA_SPARSE_SM90 is 36-70% slower than FLASH_ATTN_MLA_SPARSE
- [Bugfix] Support in-place weight updates for FlashInfer TRTLLM BF16 MoE
- [Misc] Show resolved KV cache dtype in startup logs
- [RFC]: Streaming prompt prefill for overlapping upstream generation and downstream prefill
- [Security] Validate NixlPush remote prefill params before registration
- [Security] Validate MoRI-IO remote_dp_size before EngineCore use
- Docs
- Python not yet supported