sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [JIT] Fuse FP8 KV-cache quantization into the prefix-valid commit kernel
- [AMD][Bugfix] Route default aiter NEXTN draft extend off the faulting CK kernel
- fix(inkling): return 400 instead of 500 for unfetchable/undecodable multimodal input
- [GLM47] Reduce full-assistant grammar overhead: batched masks, thinking fast path, shared metadata
- [Bugfix] Detokenizer: keep sent_offset monotonic when a character is split across steps
- [Bugfix] Skip post-experts EP all-reduce on the FlashInfer cutlass FP4 all-gather MoE path
- [AMD] Small-M W8A8 FP8 projection GEMM for Qwen3.5 AttnFP8 on gfx950
- [model] HunYuan v1: restore q/k shape after per-head qk_norm
- [qwen 3.8 next] Fuse NEXTN verify and draft graph input preparation
- [Example] Serve Jev's /v1/systemone typed-decision API on top of /v1/score
- Docs
- Python not yet supported