sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Draft] Add per-token NVFP4 MoE support for ReLU2 activation
- [Bugfix] Honor preferred sampling parameters in chat completions
- [Kernel] Reject unsupported FlashAttention versions
- fix(server): preserve SIGQUIT cleanup with multiple tokenizer workers
- [Performance] Speed up streaming function-call common-prefix detection
- fix(cache): don't stamp a Mamba checkpoint on a shorter insert key
- [FlashInfer v0.7.1] Add opt-in NVFP4 W4A16 support to FlashInfer MegaMoE
- [Speculative] Add RACER: reuse target logits for training-free tree drafting
- fix(diffusion): allow SageAttention for FLUX.2
- [Feature] Fuse DeepSeek MoE reduction, all-reduce, and RMSNorm with FlashInfer
- Docs
- Python not yet supported