sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [AMD][gfx1250] Enable fp8 wo_a bpreshuffle GEMM for MI455
- Improve total prefix cache tokens available when using hicache + write back
- [AMD][gfx95] GLM-5.3-Flash MXFP4 MoE decode: small-M kernel and sort-fused FP4 quant
- [MLX] Release finished request state between prefill-only batches
- fix(moe): Nemotron-H fix, honor is_gated=False and support relu2 in native MoE
- [diffusion] LTX-2: overlap MP4 encode with VAE decode (follow-up to #41819)
- [DSA] Reject backends that cannot serve index_kpool > 1 at arg resolution
- [AMD][Perf] Quantize the BF16 shared expert of Quark MXFP4 checkpoints to MXFP4 while loading for shared-expert fusion
- Serve decision model checkpoints natively on /v1/decisions, /v1/jev, and /v1/systemone
- [Bug] `/v1/messages`: a `stop_sequences` hit is reported as `end_turn` with `stop_sequence: null`
- Docs
- Python not yet supported