sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- ROCm: add Triton fallback for DSV4 indexer query
- [Bug] [MLX] test_batched_decode_matches_solo asserts bitwise solo/batched token equality, which no batch-shape-varying backend can provide
- [Qwen3.5] Fuse gated RMSNorm and FP8 quantization
- [MHA] Support HND KV cache layout in CPU offload
- [scheduler] Fix prefill delayer timeout rank divergence
- Free draft's duplicate embed/lm_head before KV cache sizing
- [KDA] Fuse prefill QKV causal convolution and layout
- [SM120] Add optional b12x PCIe all-reduce (one-shot + fused RMSNorm + FP8 DMA ring)
- [FA4] SM100 relative-bias decode kernel + paged dispatch
- [Bugfix] Download metadata for object-store draft models
- Docs
- Python not yet supported