sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [MiniMax-M3] Fuse SwiGLU and DeepGEMM MXFP8 quantization
- [NPU] Support Qwen4-Exp PLE CPU offload on Ascend
- Retracted input_embeds requests silently splice two generations into one successful response (send_token_offset survives the discard of output_ids)
- benchmark serving tool does not control or record radix cache state; CI silently runs a different cache protocol than users copying the same command
- [Bug] GLM-5.3-Flash vision silently broken on main: pinned transformers==5.12.1 lacks glm5_next, AutoProcessor degrades to TokenizersBackend
- [Bug] Qwen3CoderDetector: a duplicated `<parameter=NAME>` tag mid-value truncates array/object arguments and silently overwrites the earlier match
- [Bug] OpenAI-compatible API: Python and Rust render different prompts for identical POST /v1/chat/completions requests
- [Fix] Preserve ownership in deep_gemm_wrapper.transform_sf_into_required_layout (FP8 scale use-after-free)
- [Diffusion][XPU] Enable the packed Ulysses QKV all-to-all on XPU
- [Bug] HiCache staged write-back: the 128 KiB batch path passes registered host VAs to cudaMemcpyBatchAsync and faults where CanUseHostPointerForRegisteredMem == 0
- Docs
- Python not yet supported