sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- fix: hold strong references to breakable CUDA graph break inputs
- [spec] Keep DSPARK draft tokens within the target vocabulary
- [ROCm] IndexerKPool calls aiter fp8_mqa_logits unbounded; >2 GiB logits abort every TP rank (#36960 does not cover this class)
- Gateway: a worker whose metadata discovery failed is permanently registered as model_id "unknown" under IGW
- fix(gateway): don't register a worker with no model identity under IGW
- [Fix] Pad hybrid target verify to complete request groups
- [Bug] Prefill breakable CUDA graph reuses weak-ref'd break inputs across buckets → wrong greedy output / IMA; e2e validation of #37448
- [Diffusion] Accumulate chunked LoRA merges with addmm
- [diffusion] perf: FastH3 8-Step V2 native support, native Blackwell VSA, batched H3 VAE decode
- [Bug] GLM-5.3-Flash: CUDA OOM in fp8_mqa_logits during long-context prefill kills all TP ranks
- Docs
- Python not yet supported