sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- Fix stale batch fullness after chunked prefill
- [XPU] Bind sgl_per_tensor_quant_fp8/sgl_per_token_quant_fp8 on XPU
- Fix output_dtype typo in block-wise quantization docstrings
- [Fix] : support MiniMax-M3 pipeline parallelism in PD disaggregation
- [diffusion] Add support for HunyuanImage-3.0-Instruct
- `--api-key`: allow reading the value from an environment variable (avoid exposing it in the process command line)
- fix(server_args): reject speculative decoding with torch_native attention backend
- [Fix] GLM detectors: streamed tool-call arguments disagree with non-streaming
- [Fix] step3: parameterless tool calls are dropped, batched calls get mixed up
- fix(modelopt): expand is_layer_excluded for fused and model.-prefixed names
- Docs
- Python not yet supported