sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [AMD] Support compressed-tensors W4A16 linear layers on ROCm
- Test CuTe DSL W4A16 raw weight reloads and CUDA graph scales
- [AMD] Qwen3-Next: fused TP4 all-reduce + Gemma RMSNorm + per-group FP8 quant on gfx950
- [LoRA] Synchronize streamed buffer updates before publish
- [Speculative] DSpark serving: per-step host-overhead cuts, hybrid draft-pool KV injection, strip-layout conv intermediates, NIXL PD geometry guards
- [Bugfix] Fix incorrect L3 prefetch transfer rate in sglang-simulator
- [Bugfix] Honor per-module quantization settings in Phi
- [RFC] Contention-aware batching for dynamic EPLB expert migration
- [Bug] DeepSeek-V4.1 fp8 wo_a absorb GEMM is silently ~25% wrong when DEEPGEMM_SCALE_UE8M0 is false (SM121)
- [PD][MORI] Transfer DCP-replicated DSPARK draft KV
- Docs
- Python not yet supported