sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- make mxfp8 kv_cache compatability check platform aware
- [HiCache] Degrade cudaHostRegister chunk size on failure and honor chunk limit for all pools
- [AMD] Kimi-K3 - row-parallel LatentMoE tail
- fix(function_call): don't return DeepSeek DSML tool markup as content
- [Bug] Cache request finalization is tied to cache_finished_req: aborted LMCache sessions can leak
- [NPU] Gate arch35 block-FP8 -> MXFP8 requantization behind SGLANG_NPU_ARCH35_REQUANT_BLOCK_FP8
- [Bug][mem_cache] Split abort cleanup so FlexKV cancels before STORE and LMCache ends the session after it
- [Kernel] Optimize triton decoding kernels for long context
- [Feature] Automatically truncate when the maximum tokens are exceeded instead of throwing an error
- [Feature] Support varied input formats for remaining VLM
- Docs
- Python not yet supported