vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported43 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- Rewrite worker block tables after SWA eviction
- [Performance] Optimize FlatLogprobs tail slices
- [Model][Perf] Overlap Qwen3.5 GDN input projections (in_proj_qkvz / in_proj_ba) on dual CUDA streams
- [Feature]: Add sink to MLA attention
- [KV Connector] Eager KV prefetch at request enqueue time in `LMCacheMPConnector`
- [MoE Refactor] Mk construct
- docs: add Agent Friendly badge
- [Bug]: FlashInfer JIT compilation fails with "No such file or directory" in v0.20.1/v0.20.2 (docker)
- [Core][Feat] Pluggable KVCacheConfigBuilder for platform/model-specific KV cache planning
- [Bugfix][GGUF] Fix deepseek2 architecture not supported when loading without config.json
- Docs
- Python not yet supported