sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [TRT-LLM MLA] Run chunked-KV prefill under breakable CUDA graphs
- [TRT-LLM MLA] Fuse cached-prefix K/V packing and FP8 quantization
- [Simulator] Fix Mamba sizing with explicit max-total-tokens
- [Simulator] Resolve sgl_kernel CPU ops at import time under the CPU engine
- [FP8 KV] Saturate bf16->e4m3 KV/QKV casts to ±448 instead of overflowing to NaN
- [KDA] Zero every CUDA-graph padded request's output rows in the CuTe DSL MTP kernel
- [Sampling] Reject logit_bias / repetition_penalty values that overflow the sampler
- [Tokenizer] Validate caller token IDs against the vocab and active multimodal markers
- [DFlash] Feed committed tokens to the penalizers (repetition/frequency/presence/min_new_tokens)
- [MM] MultimodalInputs.merge: adopt the M-RoPE delta and clear its decode cache
- Docs
- Python not yet supported