sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- fix: guard paged allocator before kernel launch
- [Bug] DeepSeek-V4-Pro TP16 weight loading fails with out-of-range w2/w13 shard offsets in v0.5.16
- [MLX] Disable multimodal on the MLX backend (text-only inference path)
- [NPU] Support Jet-Nemotron-2B on Ascend
- [Gemma-3] RMSNorm: higher-rank q_norm/k_norm skip the fused CUDA kernel, and mixed-dtype weights return NaNs
- fix(server_args): correct false "no_buffer is the default" claim for --enable-linear-replayssm
- [Bug] v0.5.16 glm GLM-5.2 W4AFP8 + EAGLE + TP8 , MTP seed issue
- [Bug] Kimi-K3: repeated 'CUDA error: unspecified launch failure' in decode on the 07-29 kimi-k3 image (c6ad1f26), not on 74968e5653
- [Feature] Integrate NCCL RAS runtime diagnostics into SGLang health checks
- [GPU] enable pdmux on SM100 and SM120 GPUs
- Docs
- Python not yet supported