sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [MiniMax-M3] Add native SM90 Q8KV8 sparse prefill Step 3
- [Do not merge - need to clean] GLM-5.3-Flash (glm5_next) LoRA serving + glm53next port rebased
- [Intra-node PD][DSV4] Layer batching and status polling optimization
- [Mamba] Reset the ReplaySSM ring cursor on the extra_buffer donate path
- [NVFP4 MoE Marlin fallback] prepare_moe_nvfp4_layer_for_marlin leaks ~0.66 GiB/layer; old params retained by a {fqn: tensor} registry → OOM on ≤48GB cards
- [Bug] DFLASH/DSPARK draft KV pool budget uses `tp_size` instead of `attn_tp_size`, causing OOM under DP attention in Kimi-K3
- feat: add MiniCPM4.1 SparDA host-resident prefetch
- [Rust TreeCore] Torch-free RadixValue backend and suffix continuation inserts
- [Fix] MiniMax H3: honor requested inference step count
- [MoE] Add Triton FP8 compute support for NCCL EP LL
- Docs
- Python not yet supported