vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Bugfix][Frontend] Start Muse Glimmer structured outputs at to=user
- [Bugfix] Do not FULL-capture spec-decode batches in TurboQuant attention backend
- [Perf] TurboQuant: run spec-decode verify batches as decodes with FULL cudagraphs
- [Feature] Add PCP O-Proj tensor parallelism
- [Bug]: DFlash2 draft models fail to load on main — #52560 reverted the `decoder_layer_cls` indirection added by #52816
- [Bug]: Unquantized BatchedTritonExperts can change greedy top-1 across identical runs
- [Bug]: Run-to-run performance non-determinism with speculative decoding at temperature=0 (fixed seed) on DeepSeek-V4-Flash / Blackwell SM120
- [Bugfix] Replace non-deterministic index_add_ with scatter-then-sum in TopKWeightAndReduceNaiveBatched
- fix(docker): build lmcache from source in KV-connectors images to match shipped torch ABI
- [Feature]: Native support for Granite Switch (GraniteSwitchForCausalLM)
- Docs
- Python not yet supported