sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [QSA] Enable breakable prefill CUDA graphs for text-only Qwen3.8 Flash-Next (capture-safe metadata, MTP side-channel padding)
- [KDA] Zero every CUDA-graph padded request's output rows in the CuTe DSL MTP kernel
- [Sampling] Reject logit_bias / repetition_penalty values that overflow the sampler
- [DFlash] Feed committed tokens to the penalizers (repetition/frequency/presence/min_new_tokens)
- [MM] MultimodalInputs.merge: adopt the M-RoPE delta and clear its decode cache
- [HiCache] Reset decode ReplaySSM ring cursor on load-back Mamba slots
- [HiCache] Fence full prefill CUDA-graph replay behind pending load-back
- fix: route Thor SM110 FP8 and ModelOpt NVFP4 auto backends
- [HiCache][PD] Allow DSA index-K elision under HiCache and disaggregation
- [Tokenizer] Optionally omit consumed processor input_ids from tokenized multimodal requests
- Docs
- Python not yet supported