sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Bug] Multimodal CUDA IPC transport auto-selected on WSL2, where CUDA IPC is unsupported — scheduler crashes with no hint at --mm-feature-transport
- [HiSparse] HiCache as the plugin logical KV pool
- [Bug] DeepSeek-V4-Flash expert parallelism (DeepEP) crashes on Hopper/SM90 — DeepGEMM MXFP4 MegaMoE is SM100-only; late failure + misleading error
- Add NIXL as a backend for weight transfer
- feat(minimax-m3): integrate FlashInfer source MSA
- [Bug] Qwen3-VL vision features diverge from Transformers/vLLM in v0.5.17 on fine-grained grounding
- [Kernel] Optimize Triton decode attention with hardware-aware dynamic Split-K scaling (#2271)
- [Playground] Verified cell: Qwen3.8-27B / dgx-spark / nvfp4 / DFLASH2 / single
- [Bug] /health handler's timeout path doesn't cancel the scheduler-side request, causing orphaned health-check requests to pile up and crash paged-prefill batching
- fix: bypass Qwen3.5 fused CUDA preparation for mRoPE
- Docs
- Python not yet supported