sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [NPU] Fix W4A8 MoE scale_bias sharding for tensor parallelism
- [Fix] Reject Chat Completions streaming errors in bench_serving
- [HiCache] fast_file: stripe pages over several local disks and add parallel writes (write_workers)
- [Spec] Let NEXTN and the DFlash2 selector reuse a packed target lm_head
- [Spec] Build DFlash draft layers under their checkpoint names and refuse tensors the draft would drop
- [MiniMax-M3] Fix PP tc_piecewise position padding
- [MiniMax-M3][PD] Support heterogeneous TP state transfer with Mooncake
- [Fix] Request exact label logprobs for Qwen3-VL reranking
- fix(mimo-vl): keep multimodal features when capturing aux hidden states, and correct vision preprocessing
- [NPU] Set Qwen3.6 accuracy cases to greedy decoding and tune retries
- Docs
- Python not yet supported