sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- fix: warn when --dtype auto guesses float16 for a config with no declared dtype
- [Fix] Return 400 instead of 500 when input_embeds is empty
- [Bug] Qwen3.8-Flash-Next W4A16 on A100 TP4: Marlin MoE invalid thread config; Triton MoE then hits QSA FA-CuTe capture failure
- fix(qsa): support EAGLE draft-extend row bounds
- [Bug] Qwen3.8-Flash-Next QSA + NEXTN decode graph silently corrupts output on GB10 TP2
- [Bug] Spec V2 paths no longer emit speculative-decoding OpenTelemetry spans
- Capture allocator history to investigate CUDA-graph memory corruption
- AMD MI308X SGLang GLM-5.3-Flash ValueError: The checkpoint you are trying to load has model type `glm5_next` but Transformers does not recognize this architecture.
- [HiCache] Make LayerSplit MLA cache backup CP-aware
- [Kernel] Fix cross-iteration s_beta WAR race in Blackwell KDA prefill
- Docs
- Python not yet supported