sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Feature] Optimize Per-Tensor FP8 GEMM on SM120
- Unified Full KV Cache Layout in L3
- [Bug] Multi-output diffusion rollout: per-sample trajectories collapse to output 0, grouped forward AttributeError, provided latents skip packing
- [Tracking] PD disaggregation shared-protocol unification
- [MiniMax-M3] Gather-transpose batch KV pages for MSA sparse prefill (Hkv > 1)
- [Diffusion] Avoid H2D/A2A contention during layerwise offload
- [multimodal] Bound tokenizer-side GPU memory in Kimi GPU image preprocessing
- [Bug] Flashinfer backend is not supported on Blackwell GPUs
- [core] Root the JIT kernel cache under SGLANG_CACHE_DIR
- [Fix] Forward an explicit CXX to nvcc via -ccbin in JIT builds
- Docs
- Python not yet supported