sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Spec] Add --speculative-draft-kv-ratio for all-SWA DFLASH drafts
- [ROCm][Perf] Route 512-expert softmax MoE routing through aiter topk_gating
- [Bugfix] Fix route Ascend FuseEP through DeepEP path in MiniMax-M3
- [ROCm] Fuse GLM-5.2 indexer preparation on gfx950
- fix(parser): keep trailing visible text in qwen3_coder one-shot parsing
- [Bug] LoRA adapters with use_rslora=True are served with the wrong scale
- [AMD][Bugfix] Route default aiter NEXTN draft extend off the faulting CK kernel
- [NPU] Fix W4A8 MoE scale_bias sharding for tensor parallelism
- [Spec] Let NEXTN and the DFlash2 selector reuse a packed target lm_head
- [Spec] Build DFlash draft layers under their checkpoint names and refuse tensors the draft would drop
- Docs
- Python not yet supported