sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [MM] Add opt-in exact multimodal embedding row-count check at final placement
- [Fix] strip_thinking_cache: never free KV the radix tree owns after a re-prefill
- [HiCache][PD] Allow DSA index-K elision under HiCache and disaggregation
- [Tokenizer] Optionally omit consumed processor input_ids from tokenized multimodal requests
- [PD][Spec] Give a PD decode's first EAGLE draft step a proposal distribution
- [DSA] Q8KV8 FP8 Sparse Prefill on GLM-5.3-Flash with BF16 KV Cache
- [AMD][Quark] Kimi K3 Quark FP8/MXFP4 fusion
- fix(function_call): process all buffered Apertus tool blocks on stream end
- fix(function_call): recognize abbreviated Harmony commentary header in gpt-oss detector
- [Bug] PEFT adapters with bias="lora_only" or "all" crash or silently corrupt weights during normalize_qkv_proj
- Docs
- Python not yet supported