vllm
https://github.com/vllm-project/vllm
Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported51 Subscribers
View all SubscribersAdd a CodeTriage badge to vllm
Help out
- Issues
- [CI][Rust Frontend] Add Dockerfile for mock engine
- [Perf] Add exact distributed MRV2 prompt logprobs
- [Core][KVConnector] Async CPU KV-offload initialization
- [Bugfix] avoid IndexError for trailing dotted CLI args
- [Bugfix] create the weight lock directory itself
- [Docs] Fill Apache LICENSE copyright holder
- [Performance]: gfx1100 ranks ROCM_ATTN above TRITON_ATTN for a custom kernel it does not use at head_size 256
- [ROCm] Rank ROCM_ATTN by whether its custom kernel applies on RDNA
- [Bugfix] Don't let a structured-output request sample from an unmasked row
- [Bugfix][Quantization] Reject NVFP4 checkpoints with missing global scales (linear + MoE)
- Docs
- Python not yet supported