sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Bug] LMCache and FlexKV ignore the effective committed KV length when storing finished requests
- [AMD] Fix Quark W4A4 MXFP4 MoE clamped-SwiGLU on the AITER path
- fix: honor effective committed KV length in external caches
- [AMD] [Bugfix] Stop registering oversized inputs in the AMD deterministic all-reduce
- [MiniMax-M3] Add one-page sparse attention paths
- [Cherry-pick to release/v0.5.17] Backport 2-CTA TGV exit barrier (#32954)
- feat(serving): add --enable-sort-tool-schema-keys to share prefix cache across tool key orders
- [Fix] Correct inverted assertion message for enable_mixed_chunk with speculative decoding
- [ROCm][Bugfix] Kimi-K3: opt-in gfx942 MXFP4-to-int4 MoE conversion
- [PD] PP prefill: fold HiCache prefetch readiness into the bootstrap consensus
- Docs
- Python not yet supported