sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Bug] DP-attention scheduler crashes (NoneType len) when attn_tp_size>1 and attn_cp_size>1: request-broadcast source rank has work_reqs=None
- [Mamba] Warn when prefix-cache checkpoint capacity is insufficient
- fix(glm4v): align mRoPE mask with final input IDs
- [2/N] Quantization Refactor: drive the method registry from a string spec table
- Retry queued prefills when Mamba cache capacity may have recovered
- feat(hicache): reserved overflow ring for mamba host pool write-backup
- fix(hisparse): 4-18x SM120 planned-copy collapse at non-64B strides
- fix(tokenizer): keep generic TokenizersBackend when re-resolve fails (GLM-5.3)
- [unified-memory] Speculative decoding on the unified memory pool: translate every read and write path
- [unified-memory] Fuse the draft model's KV into the target's page envelope (mechanism)
- Docs
- Python not yet supported