sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Bug] [ROCm] DeepSeek-V4.1 FP4 indexer still exhausts HBM under long-context AgentX after #37660
- [XPU] Add HiCache L1/L2 transfer backend
- [Bug] DSA KPool decode JIT specializes on exact block-table width
- [Fix] Make the Marlin atomicAdd reduction opt-in (SGLANG_MARLIN_USE_ATOMIC_ADD) and force it off under deterministic inference
- Add --padded-output-tokens to close the completion-length oracle
- errors on model type `qwen4_exp` for nvidia/Qwen3.8-Flash-Next-NVFP4
- [diffusion] Support MiniMax-H3 PDD Acc LoRA (8-NFE)
- [Bug] NEXTN on B200 silently switches Qwen3.5-122B to triton when --page-size is set; forcing trtllm_mha is 2.6x faster at long context
- [model-gateway] Add cache-aware policy decision metrics
- [Bug] Kimi K3 leaks raw XTML tool calls into content when no tools are declared
- Docs
- Python not yet supported