sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- Prefill Context Parallel Support for Qwen3.8-Next-Flash
- [AMD] Reuse AITER DSA prefill metadata across layers
- [Bug] Quantized DFlash2 draft silently yields ~0% acceptance — no error, no warning (the quiet counterpart to #36599)
- [DSV4.1] Enable Hopper decode optimizations for TP8
- [Bug / Security] Potential DFA state explosion and CPU thread hanging in JSON Schema grammar compilation under deeply nested schemas
- Fix hybrid session references after an empty terminal cache insert
- [HiCache] Preserve L2/L3 hit attribution across scheduling retries
- MoE deferred finalize is unreachable for models that supply their own routing (FLASHINFER_TRTLLM_ROUTED excluded)
- [Bug] GLM-5.x NoPE MLA (qk_rope_head_dim=0) cannot run on SM120: every DSA sparse-MLA backend is unavailable
- [DSA][DCP] Decode context parallel for the DSA attention family (GLM-5.3 / DeepSeek-V3.2) on Hopper
- Docs
- Python not yet supported