vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- fix: resolve issue 57532 by passing quant_config=None to SharedHead
- [Frontend][Core] Sampled-token logprob fast path for /inference/v1/generate (return_token_logprobs)
- [Bugfix] Restore Aria missing-weight validation
- fix: return loaded parameters in Aria load_weights
- [Bug]: On SM12x, fp8 block linear still selects DeepGEMM when E8M0 is disabled, which now hard-fails after the a6bbb80 pin
- [Installation]: [ROCm][gfx1151] q_gemm.cu fails to compile: missing half and half2 atomicAdd overloads
- Fix: Internlm tool parser loses tool calls on chunked streams
- [Bugfix] Fail closed in is_uva_available() when the UVA alias is not live
- [Kernel] Add tuned W8A8 block-FP8 GEMM configs for gfx950 (MI350X)
- [Bug]: [ROCm][gfx1151] ROCM_ATTN returns different outputs for the same greedy request after other requests
- Docs
- Python not yet supported