sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [AMD] Enable aiter MLA for 12-head models via 12->16 zero-pad (Kimi-K3)
- fix(kimi-k3): mint opaque unique tool_call ids instead of positional name:ordinal
- [Attention][DCP] Add symm_a2a: single-node symmetric-memory A2A backend for MLA decode
- [sgl-router]: Fix sgl-router Docker port and CLI usage
- feat: 1/N DeepSeek V4 CP Cache LayerSplit: Common Infrastructure
- Support gated launch to defer startup memory allocation
- [Bug] Unlimited-OCR: server crashes when batching images with different gundam tile counts
- fix(dsv4): preserve BF16 precision in KV cache
- [AMD] Optimize MLA for Kimi K3, ~70% imporvement in MLA for K3 @ fp8 KV
- Fix #19612: Unified JIT and Precompilation Cache Directory
- Docs
- Python not yet supported