sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Bug] Qwen3.8-Flash-Next QSA + NEXTN decode graph silently corrupts output on GB10 TP2
- [Bug] Spec V2 paths no longer emit speculative-decoding OpenTelemetry spans
- Capture allocator history to investigate CUDA-graph memory corruption
- AMD MI308X SGLang GLM-5.3-Flash ValueError: The checkpoint you are trying to load has model type `glm5_next` but Transformers does not recognize this architecture.
- [HiCache] Make LayerSplit MLA cache backup CP-aware
- [Kernel] Fix cross-iteration s_beta WAR race in Blackwell KDA prefill
- [Perf] DeepSeek-V4-Flash-0731 on 4×H800 (SM90, TP4, no EP): ~40% lower per-stream decode at conc=32; dp-attention regresses and TBO is gated — what is the recommended small-node config?
- [Model] Qwen3.5-MoE: RuntimeError when using --quantization-param-path for FP8 KV cache
- [Bug] Kimi-K3 chunked prefill submits different VocabParallelEmbedding ALLREDUCE sizes across TP ranks
- [CI][Diffusion] Add caching for npu GT images in CI
- Docs
- Python not yet supported