sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Feature] Serve Qwen3.8-Flash-Next (176B / 6B-active MoE) at 2.57 bpw on one 24 GB GPU with 32 GB host RAM
- fix(qwen4): CPU-offload correctness, breakable graphs, mmap PLE table
- fix(spec): commit PLE state after ReplaySSM verify, NGRAM on Qwen4-Exp
- feat(moe_wna16): 2-bit experts, expert streaming, N-contiguous GEMV
- feat(mem_cache): elastic VMM expert row arenas and lazy KV backing
- feat(qsa): quantized KV pools with dequant-on-gather (fp8/int8/int4)
- [NPU] Support Kimi-K3 decode context parallelism
- fix(inkling): prevent server crash from grouped GEMM on SM120
- [ci] Support batched multi-node NPU e2e cases in full-test-npu
- [diffusion] Calibrate residency probes that run a single iteration and honor pipeline step minimums
- Docs
- Python not yet supported