sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [AMD][CI] Add michaelzhang-ai to .github CODEOWNERS
- [Perf] Fuse H64 FlashMLA prefill statistics conversion
- [Bugfix] Strip repeated leading think markers in streaming reasoning parser
- [Feature] Rust mm datapath: InternVL server-pipeline processor
- Reject unsupported trained bias in LoRA adapters
- [Bug] --cuda-graph-max-bs-prefill silently rounds down to the nearest capture bucket (decode does not), which can disable post-capture KV sizing
- Benchmark: Qwen3.8-27B FP8/BF16 on H200 with float32 SSM state (46 speed rows, 6 full GSM8K scores)
- [Bug] GLM-5.3-Flash NVFP4 at TP4 loops in reasoning with no final answer on B200/B300
- [diffusion][xpu] FLUX.2: bit-exact Triton LayerNorm+modulate and gated residual norm for XPU
- [diffusion] FLUX.2: build the single-block output projection input without a cat
- Docs
- Python not yet supported