sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- fix(server): reject torch native speculative decoding
- Reject torch_native attention backend + speculative decoding at startup
- Register SGLang kernels in Pytroch dispatcher
- [Bugfix] Return the longest partial in _ends_with_partial_token
- [Feature] Generic OCP MXFP4 KV cache support
- fix(metrics): handle missing CPU sequence lengths
- [CI] Stop run_suite from stalling after every result is reported
- bench_one_batch silently ignores --moe-runner-backend (initialize_moe_config is scheduler-only): quantized MoE benches run the BF16/triton fallback
- [diffusion] bench: report steady-state denoise latency by default
- [Docs] Prefer FlashInfer Mamba for Nemotron 3 Ultra
- Docs
- Python not yet supported