vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported36 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix][ROCm][CI] Restore the DeepSeek-V4 input GEMM override point
- [Bug]: deepseek-v4-flash-0731 use vllm 0.27.0 run error
- fix(pooling): register BGE-M3 combined task processor
- [Bugfix] vLLM crashes at startup when DeepEP v2 is used with `--enforce-eager` wiht TRTLLM Bf16
- [Bug]: Optional usage telemetry (py-cpuinfo) can kill engine startup in a forked worker
- [3/N] Harden Transformers modelling backend multi-modal path
- [Bug]: Gemma-4 NVFP4 (fp8 KV) on Blackwell: FLASH_ATTN is selected, silently falls back to FA2, then fails with "FlashAttention only support fp16 and bf16 data type"
- [Bug]: DeepSeek V4 Python renderer misplaces tools when a system message exists
- [CI] Report torch-nightly results to PyTorch CRCR
- [Model] Support R3 capture with DeepGEMM MegaMoE
- Docs
- Python not yet supported