vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Doc] Update extension names in incremental build guide
- [ROCm][Quantization] Match Quark exclude semantics in the shared-expert FSE check
- [Bug] DeepSeek-V4-Flash: non-deterministic output at temperature=0, rate scales with concurrency
- [Bug]: Scheduler permanently stops admitting requests once running + skipped_waiting ("deferred") reaches max_num_seqs - engine reports healthy, KV cache mostly free, only a restart recovers
- [Bug]: DC unavailable for GlmMoeDsa (GLM-5.2 FP8) on SM90 — sparse MLA backend lacks decode-LSE support
- vLLM v0.26.0 + Ray TP8: intermittent wedge, `RayWorkerProc rank=[0] died unexpectedly` after `shm_broadcast` block starvation; API stays up but generation is dead
- [Bugfix][Frontend] Support Prometheus name[] metric filtering
- [7/N][KV Connector][NIXL] Stage host KV reads through device memory
- [Model] Enable LoRA support for tower and connector in GLM-ASR
- [Mamba] Support quantized FlashInfer ReplaySSM state cache
- Docs
- Python not yet supported