sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [NPU] [Bugfix] Fix rope sin cos pre-calculation for Qwen3.5 and Qwen3_next during Pipeline Parallelism
- [NPU] Support chunked prefill for DeepSeek-V4
- [HiCache] Unified Mooncake Registration for Logical Anchors and Draft Pools
- [Bug] DeepEP low_latency buffer lazy init fails during CUDA graph capture with PP=2, TP=8, DP-attention, EP=8 on Kimi K2.6 W4A8
- refactor: make customized_info msgpack-native (drop PickleWrapper)
- docs: fix README news date
- [DO NOT MERGE][AMD] Bump Mooncake pin to include cross-node RDMA multi-protocol fix
- Avoid blocking scheduler on health check send
- [Fix] Avoid chunked-prefill slot double-count that pins the running batch below `max_running_requests`
- [NPU] FIX CMB illusion of garbled characters acc problems, in prefix cache mtp scenarios.
- Docs
- Python not yet supported