sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [AMD][DSV4] perf: fold the padded-weight zeroing into the fused append kernel
- [HiSparse] Implement get_cpu_copy/load_cpu_copy for HiSparseDSATokenToKVPool
- [Bug] NVFP4 MoE + EP>1: globally-loaded input scales not sliced to local experts in non-CuteDSL branch (_compute_gemm1_alphas shape mismatch)
- [Bug] Disabling DSA via index_topk override with index_share_for_mtp_iteration=true asserts in scheduler (get_dsa_index_topk)
- [Bug] [CPU] topk_sigmoid_cpu reads past small expert buffers with correction_bias
- [Bug] qwen3.8 flash next rocm bug
- [CPU] FP8 KV store: keep fused decode path, stop using Python index_put
- [Bug] NEXTN/MTP speculative decoding fails to load MTP MoE weights under TP>1 for Glm5NextForConditionalGeneration (GLM-5.3-Flash)
- [Bug]
- [Spec] Support ReplaySSM spec-verify for DFlash and DSPARK on hybrid GDN models
- Docs
- Python not yet supported