sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Bug] SGLANG_SIMULATE_ACC_LEN silently degrades detokenization to O(n²) — `predict.fill_(100)` emits a byte-fallback token, and the `endswith("\ufffd")` commit gate then never advances the incremental-detokenization offsets
- [Bug][Diffusion] Native-fallback component loading drops all CPU-offload decisions (should_offload called without model_config) → fatal OOM on 8GB GPUs
- [Bug] TypeError in set_mamba_track_indices_from_reqs during NEXTN TARGET_VERIFY — mamba_next_track_idx is None (hybrid-mamba + speculative decoding + lazy buffer mode)
- [Bug] Kimi K3 decode crash: DSPARK target_verify + DCP → cumsum(extend_prefix_lens=None) in layers/dcp/planner.py
- fix(disaggregation): route equal-TP fragmented KV transfer through staging
- [Bug] Multimodal CUDA IPC transport auto-selected on WSL2, where CUDA IPC is unsupported — scheduler crashes with no hint at --mm-feature-transport
- [HiSparse] HiCache as the plugin logical KV pool
- [Bug] DeepSeek-V4-Flash expert parallelism (DeepEP) crashes on Hopper/SM90 — DeepGEMM MXFP4 MegaMoE is SM100-only; late failure + misleading error
- Add NIXL as a backend for weight transfer
- feat(minimax-m3): integrate FlashInfer source MSA
- Docs
- Python not yet supported