sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Spec][MegaMoE] Let the speculative draft choose its own W4A4 MXFP4 MegaMoE MMA type
- [CP 5/5] Test GLM-5.3-Flash CP4 with KDA and MoE TP4 on B200
- [Spec] LiLiCorr: named head MLPs and quantized head linears
- [PD] Fix crash for allocation when mamba + pd + decode radix cache
- [NPU] [quantization] support NZ for w4a4 linear
- [QSA] Enable breakable prefill CUDA graphs for text-only Qwen3.8 Flash-Next (capture-safe metadata, MTP side-channel padding)
- [KDA] Zero every CUDA-graph padded request's output rows in the CuTe DSL MTP kernel
- [Sampling] Reject logit_bias / repetition_penalty values that overflow the sampler
- [DFlash] Feed committed tokens to the penalizers (repetition/frequency/presence/min_new_tokens)
- [MM] MultimodalInputs.merge: adopt the M-RoPE delta and clear its decode cache
- Docs
- Python not yet supported