sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [AMD] Qwen3-Next: fused QK Gemma RMSNorm + partial RoPE on gfx950
- [mem_cache] Rename `cache_unfinished_req` to `checkpoint_req` and count cache hits only when a request finishes
- [Bug] Remove unused Req.decoded_text and dead stop-string fallback
- [XPU] Allow the intel_xpu DSA backend with HiSparse
- [CPU] add chunk_kda kernel for xeon
- [diffusion] fix: fall back from SageAttention on unsupported head sizes
- [AMD] Fix HiSparse slots in fused BMM KV writes
- [Intel GPU] Add support for MXFP8 W8A8 compressed-tensors scheme
- [ROCm] Use the fused MLA absorb + RoPE + KV-write kernel for decode-sized forward modes only
- [AMD] Add opt-in radix-select top-p pivot to the ROCm Triton renorm
- Docs
- Python not yet supported