vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported49 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [Kernel][ROCm] W8A8 FP8 hip kernel for RDNA4
- [Mypy] Fix mypy typing for Ultravox and Unlimited-OCR models
- [Bug]: cutlass_scaled_mm_supports_block_fp8 excludes SM89 (Ada Lovelace) despite native FP8 tensor core support
- [Mypy] Fix mypy typing for Voxtral and vision models
- [Mypy] Fix mypy typing for Whisper models
- [Mypy] Fix mypy typing for Zamba2 models
- [ROCm][Bugfix][Spec Decode] Add compact_topk_indices to the ROCm DSA MTP drafter
- [CPU][Whisper] Support W4A16 quantized Whisper on the CPU WNA16 kernel
- [Bugfix][Whisper] Add HF generate_with_fallback on LLM.generate
- fix(logging): keep structured JSON log lines parseable
- Docs
- Python not yet supported