sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- errors on model type `qwen4_exp` for nvidia/Qwen3.8-Flash-Next-NVFP4
- [diffusion] Support MiniMax-H3 PDD Acc LoRA (8-NFE)
- [Bug] NEXTN on B200 silently switches Qwen3.5-122B to triton when --page-size is set; forcing trtllm_mha is 2.6x faster at long context
- [model-gateway] Add cache-aware policy decision metrics
- [Bug] Kimi K3 leaks raw XTML tool calls into content when no tools are declared
- [AMD] Add FlyDSL GDN prefill backend
- [Bug] Match OpenAI tool schema field order and omission semantics for `strict` and `defer_loading`
- [feat] host-staged all-reduce for two GPUs without peer access
- [Tracing] Merge async trace flushes across scheduler batches
- [Tracking] Pipeline parallelism x speculative decoding (EAGLE/MTP): converge prefill gate PRs, make the path model-agnostic, decide decode design
- Docs
- Python not yet supported