vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Feature] Add PCP O-Proj tensor parallelism
- [Bug]: DFlash2 draft models fail to load on main — #52560 reverted the `decoder_layer_cls` indirection added by #52816
- [Bug]: Unquantized BatchedTritonExperts can change greedy top-1 across identical runs
- [Bug]: Run-to-run performance non-determinism with speculative decoding at temperature=0 (fixed seed) on DeepSeek-V4-Flash / Blackwell SM120
- [Bugfix] Replace non-deterministic index_add_ with scatter-then-sum in TopKWeightAndReduceNaiveBatched
- fix(docker): build lmcache from source in KV-connectors images to match shipped torch ABI
- [Feature]: Native support for Granite Switch (GraniteSwitchForCausalLM)
- [Spec Decode] Add a static regression guard for DFlash2 decoder_layer_cls
- [CI] Persist AMD test caches and defer collection-time downloads
- [Platform] Allow pinning the attention backend for components that auto-select
- Docs
- Python not yet supported