sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [TRT-LLM MLA] Fuse cached-prefix K/V packing and FP8 quantization
- [FP8 KV] Saturate bf16->e4m3 KV/QKV casts to ±448 instead of overflowing to NaN
- [KDA] Zero every CUDA-graph padded request's output rows in the CuTe DSL MTP kernel
- [Sampling] Reject logit_bias / repetition_penalty values that overflow the sampler
- [Tokenizer] Validate caller token IDs against the vocab and active multimodal markers
- [DFlash] Feed committed tokens to the penalizers (repetition/frequency/presence/min_new_tokens)
- [MM] MultimodalInputs.merge: adopt the M-RoPE delta and clear its decode cache
- [HiCache] Reset decode ReplaySSM ring cursor on load-back Mamba slots
- [HiCache] Fence full prefill CUDA-graph replay behind pending load-back
- [Quant] Reject non-finite and non-per-tensor FP8 KV scales at load
- Docs
- Python not yet supported