sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- Bound eager input storage by forward capacity
- Skip unused decode graph input embeddings
- feat: add NVFP4 DSA KV cache for TRTLLM-GEN
- [DCP + L3 3/N] Support hybrid MLA/KDA DCP with file and Mooncake HiCache L3
- [Benchmark] Add random shared-prefix dataset
- [cherrypick from #36389] [sglang][lora] Support DP attention in LoRA backends
- [Session] Wait for all DP session-open acknowledgments
- [Fix] Abort over-length embedding requests before enqueuing
- [Bug] SM120: block-FP8 linears ignore scale_fmt "ue8m0" for activations (cutlass path uses FP32 amax/448 scales)
- [AMD] [GLM5] Use Triton sparse attention for GLM-5.3 on gfx950
- Docs
- Python not yet supported