sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Feature] Disable multimodal weights loading to free GPU memory
- fix: correct MHA PP KV-pointer slicing for uneven full-attn splits (#27740)
- Add manual streaming-session throughput benchmark
- refactor(attn): rename init_cuda_graph_state -> init_static_metadata_buffers
- fix(frontend): reject out-of-range top_logprobs_num
- [NPU] Support w8a8_int/w4a4_int online quantization
- [DSA] perf: Fuse dummy topk in only K branch
- qwen3_next: support quantized .weight_packed/.qweight in _make_packed_weight_loader
- [Bugfix] Preserve single quotes inside JSON values in Llama32 streaming tool calls
- [HiSparse] Optimize Quest sparse page selection (Part 2 of 4)
- Docs
- Python not yet supported