vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported43 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix] v1/sample: use out-of-place scatter in apply_top_k_top_p_pytorch
- entrypoints/openai: skip tool parser in streaming when tool_choice="none"
- [Bug]: Forked workers retain stale CUDA primary contexts from parent process
- [Perf] Eliminate two GPU→CPU syncs in make_kv_sharing_fast_prefill_common_attn_metadata
- Release stale CUDA primary contexts inherited by forked workers
- Enable native Windows CUDA source build
- [Bugfix] Fail fast when xxhash prefix cache hashing lacks dependency
- fix(v1): enforce VLLM_ALLOW_INSECURE_SERIALIZATION in run_method
- [Model][Experimental] DeepSeek V4 w4a4 MegaMoE support
- [Platform][CUDA] Warn on PTX fallback and optionally invalidate Triton cache (sm_121)
- Docs
- Python not yet supported