vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported36 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Kernel] Remove unused N/B constexpr params from pack/unpack seq kernels
- Fix MTP Mamba align prefix cache retention
- [ZenCPU] Changes for relevant cpu tests and image build for ZenDNN
- [Bug] Optional Transformers model submodules can break vLLM model imports
- [Bug]: Failed to initialize FlashInfer All Reduce workspace: CUDA driver error: invalid device ordinal
- fix(security): validate bad_words token IDs against model vocab size
- Perf/h20 moe config e256 n512
- [Bug]: enable_multithread_load silently disables EP weight filtering (local_expert_ids never reaches multi_thread_safetensors_weights_iterator)
- [Bug]: M-RoPE recomputation crashes or mis-associates media at a chunked-prefill boundary
- [ZenCPU] Changes with respect to cpu platform detection test for ZenDNN
- Docs
- Python not yet supported