sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Perf] DeepSeek-V4-Flash-0731 on 4×H800 (SM90, TP4, no EP): ~40% lower per-stream decode at conc=32; dp-attention regresses and TBO is gated — what is the recommended small-node config?
- [Model] Qwen3.5-MoE: RuntimeError when using --quantization-param-path for FP8 KV cache
- [Bug] Kimi-K3 chunked prefill submits different VocabParallelEmbedding ALLREDUCE sizes across TP ranks
- [CI][Diffusion] Add caching for npu GT images in CI
- fix: hold strong references to breakable CUDA graph break inputs
- [spec] Keep DSPARK draft tokens within the target vocabulary
- [ROCm] IndexerKPool calls aiter fp8_mqa_logits unbounded; >2 GiB logits abort every TP rank (#36960 does not cover this class)
- Gateway: a worker whose metadata discovery failed is permanently registered as model_id "unknown" under IGW
- fix(gateway): don't register a worker with no model identity under IGW
- [Fix] Pad hybrid target verify to complete request groups
- Docs
- Python not yet supported