sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [DO NOT MERGE] test mooncake-transfer-engine from test.pypi
- QSA decode has no working kernel path on SM121 (GB10 / DGX Spark); Qwen3.8-Flash-Next unservable
- [Qwen3.8] Fix FP8 KV cache support in QSA
- [Diffusion] Preserve mapped component weights during low-memory offload
- TBO unreachable for GLM-5.3-Flash (DSA+mamba hybrid): silent --enable-dp-attention reset + missing Glm5NextDecoderLayer strategy
- Name the condition that blocks flush_cache / attach / detach
- Fix TP decode poll collective ordering
- [GB10] Long prefill (>40k tokens) exhausts unified memory and silently kills the worker rank — no traceback, no OOM record; cross-stack control passes at 54k
- [Spec][Ngram] Expand SAM match length beyond max_trie_depth
- [AMD] Enable the GDN decode projection/Conv1D fusion on ROCm
- Docs
- Python not yet supported