sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [CPU] Fold the MXFP4 block scale in 2 instructions instead of 4
- [AMD][CI][Fix] Guard FP32 LM head mm(out_dtype) fast path on ROCm
- PP8 disaggregated prefill has a load-independent ~30 s TTFT floor on Kimi-K3
- [multimodal] Classify PIL decode-time failures as client errors (400)
- [Perf] No MMQ kernel for I-quant GGUF: 4-6x slower prefill than llama.cpp on the same weights (measured)
- feat(sgl-router): decompose the pre-first-token path into phase histograms
- [EPD] Support L2-only multimodal embedding cache with rank-0 allocation
- Replace vLLM operators with custom operators from the sglang kernel
- [Kernel] Barrier-free 2shot_push custom all-reduce for the mid-size band
- [sgl-kernel][CPU] Add group-aware CPU SHM collective kernels
- Docs
- Python not yet supported