sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Feature] Sparse destination-aware MoE dispatch to skip inactive EP ranks
- [Fix] Match compressed-tensors weight_packed names when loading MXFP4 MoE
- [Bug] Kimi-K3 output abnormal when high pressure
- [Bug] GLM-5.3-Flash crashes at startup under pipeline parallelism — KeyError: 'residual'
- [Bugfix] Use the canonical get_local_ip_by_remote in dumper
- fix(vlm): classify malformed lazy image decode as bad request
- fix(scheduler): advance the input-logprob offset for skipped prefill requests
- [Frontend] Linearize GLM tool-result ordering
- [Speculative] Add bounded compact DFlash KV cache for PD serving
- fix: retry post-warmup freeze_gc so --skip-server-warmup doesn't race server startup
- Docs
- Python not yet supported