sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [AMD][gfx95] GLM-5.3-Flash MXFP4 MoE decode: fold routing and FP4 quant into sort
- [Fix] scheduler: stop latching batch_is_full when nothing can run
- add SGLANG_ENABLE_CHAT_ENCODE_THREAD_POOL to enable next request preprocessing ahead.
- [NPU] Resolve the DSA KV cache layout from a startup capability
- [GLM-5.3] Map forget_gate names in compressed-tensors ignore lists
- [Ascend][NPU] Add int8 KV cache for DSA models on Ascend NPU backend
- [AMD] Opt-in AITER FlyDSL causal conv1d prefill and GDN decode for Qwen3.5/Qwen3.8Pr/gdn conv flydsl
- [AMD] AITER FlyDSL hc_mix for Qwen3.8 hyper-connections on gfx950
- [AMD] Opt-in AITER FlyDSL paged decode for Qwen3.8 QSA sparse attention
- [MLX] Release finished request state between prefill-only batches
- Docs
- Python not yet supported