sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [HiCache] Preserve L2/L3 hit attribution across scheduling retries
- MoE deferred finalize is unreachable for models that supply their own routing (FLASHINFER_TRTLLM_ROUTED excluded)
- [Bug] GLM-5.x NoPE MLA (qk_rope_head_dim=0) cannot run on SM120: every DSA sparse-MLA backend is unavailable
- [DSA][DCP] Decode context parallel for the DSA attention family (GLM-5.3 / DeepSeek-V3.2) on Hopper
- [Bug] --enable-mixed-chunk corrupts mamba radix cache checkpoints on hybrid GDN models (mixed batch skips extra_buffer write, slot still donated)
- [Bug]Unauthenticated PUT /route on Bootstrap HTTP allows route poisoning and metadata redirection
- [Bug] ep_scatter_from_psum missing expert_start/num_experts args → TypeError on deepep_v2 prefill
- [Bug] HiCache write_through does not fully persist first-seen prefixes before eviction
- [Tracking] Upstream remaining sglang-miles changes to main
- [Refactor][2/2] Rename FFN preparation APIs and runtime symbols
- Docs
- Python not yet supported