sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- Resolve pr/feature/optimize glm53 flash frontier
- Allow 128 chat logprob candidates through the model gateway
- [HiCache] Demote an internal node's SWA and Mamba state on write_back eviction when its Full KV is already on host
- [VLM] Rebuild padded input ids after caller mm_hashes
- [VLM][EPD] Fix tokens-in requests under zmq_to_scheduler
- fix flatten bucket rank device
- Create a workflow to force clean up ccache
- Fix issues when deploy modelopt Llama-3_3-Nemotron-Super-49B-v1-FP8
- Fix random token sequence length in bench_serving
- [Roadmap][Feature] Support Moore Threads (MUSA) GPU
- Docs
- Python not yet supported