vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix] Remove activation override in test_flashinfer_cutedsl_fp4_moe
- [Doc]: metrics page is missing the spec-decode counters and cache_config_info
- [Docs] Document metrics whose names are set through variables
- [Doc]: prompt_lookup_min/prompt_lookup_max docstrings give the wrong defaults
- [Doc] Fix prompt_lookup_min/max defaults in SpeculativeConfig docstrings
- [MoE] Complete transition of packed mxfp4 quantization methods to the mxfp4 oracle
- [Bugfix][Mamba] Zero mamba state blocks on (re)allocation to stop stale-state poisoning on hybrid models
- [Bugfix][Mamba] Guard state-block prefix-cache hits behind the writing step's commit fence
- [Perf][Qwen] Fuse FP8 Quant into Qwen 3.8 AllReduce
- [Bugfix] Route every speculation-capable row through the speculative path so recurrent state stays consistent
- Docs
- Python not yet supported