sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [diffusion] fix: stop silently dropping weights in update_weights_fro…
- [Feature] Optimize FP8 Blockwise GEMM on SM120
- [Bug] All schedulers hang in cuModuleLoadData on first EAGLE verify (DSA attention, PD-disagg decode), watchdog timeout
- [Feature] Support shared to sparse experts fusion for Qwen3.5 / Qwen3.6 MoE on SM120 (blockwise FP8)
- [Feature] Finish the B12X FlashInfer NVFP4 MoE integration for SM120 (#29190)
- unified_cache: MAMBA component nodes are pruned instead of downgraded on device eviction, breaking host-tier loadback
- [diffusion] MiniMax-H3 quality=high (Cache-DiT tier) topology-locked to 4xH200 — request 8xB300 audit
- [4/N] elastic-ep: Scale at DP-attention replica granularity
- [AMD/ROCm] Kimi-K3 KDA prefill occasional hang at c=1: chunk_kda_fwd → tolist() → hipMemcpy D2H never returns
- [AMD] perf: fuse MLA value projection with MXFP4 output quantization
- Docs
- Python not yet supported