sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [PDMux][Tracking] GLM-5.3-Flash and DeepSeek-V4.1-Flash integration
- [AMD] Kimi-K3: support aiter MegaMoEV2 fused EP MoE on ROCm
- [Bug] [Diffusion] Hybrid SP+TP does not work correctly for the I2V models
- [AMD][GLM-5.3-Flash] Allow shared experts fusion on gfx942 (MI308X)
- Improve total prefix cache tokens available when using hicache + write back
- [RFC][NPU] INT8 (C8) KV cache for DSA models (GLM-5.2) on Ascend A2/A3 (910B/910C)
- [AMD][Kimi-K3] Update GEMM dispatch to route the fused MoE front GEMM through Aiter
- [AMD] gfx1250: fix GDN chunked-prefill GPU fault and state corruption
- [HiCache] Resolve cudaMemcpyBatchAsync from the JIT module's own CUDA runtime
- [AMD] GLM-5.3-Flash: fuse shared expert and KDA projections on Quark MXFP4
- Docs
- Python not yet supported