sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- Allow CPU image preprocessing to preserve serving GPU memory
- Report which prefill phases consume GPU memory before an OOM
- Warm up real serving paths before the first user request
- Retry queued prefills when Mamba cache capacity may have recovered
- [Bug] --enable-linear-replayssm forces no_buffer, which degrades mamba prefix caching and inflates TTFT up to 4.7x
- [diffusion] Plan component residency from calibrated warmup records (2/4)
- fix: report DeepSeek-V4 KV cache memory usage
- fix(dspark): isolate draft MoE from target expert recorder
- MLA decode-side KV broadcast: one rank pulls over the network, relays over NVLink
- feat: add fused CuTe DSL BF16 GEMM SiTU kernel
- Docs
- Python not yet supported