sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Bug] Kimi VL GPU memory usage too high
- Fix compatibility issues with ModelScope models in ServerArgs
- [Models] Add NVIDIA Eagle 2.5 VLM
- [AMD][Quantization] Online MXFP4 quantization 3/N - online `fp8` quantization per layer while loading weights, reducing memory overhead
- [Qwen3-Next] Optimize Prefill Kernel, add GDN Gluon kernel and optimize cumsum kernel
- [AMD] enable aiter attention backend for non-power 2 rope/nope models
- Fix non-determinism on Kimi-K2.5 MoE
- [Agentic Inference] Programmatic KV Cache for Agentic Workloads
- [NPU] Support mlaprolog, fp8 KVcache(only MLA) and ChunkPrefill for A5(950PR/DT) NPU
- [LoRA] Support LoRA under DP attention: idle-forward guards, attn-TP-local slicing and buffer sharding
- Docs
- Python not yet supported