sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Bug] DSV4F0731+DSPARK ERROR IN SGLANG 0.5.18
- [CPU][Diffusion] Add packed QKV kernel and enable Ulysses SP on CPU
- [weight cache] Add disk-vs-daemon parity harness; fix persistence, zombie-PID liveness, and SM100 block-FP8 gating
- fix(constrained): require output after reasoning
- [Bug] sglang-kernel JIT: HiCache L3 segfault on first backup with nvcc 12.8 + CUDA 13 runtime (compile-time vs runtime cudaMemcpyBatchAsync ABI in staged_write_back.cuh)
- [Scheduler] Add cache-hit-aware over-admission for bounded Mamba requests
- [Bug] Deterministic inference hangs when chunked prefill is smaller than the prefill alignment
- Enable and fuse the Inkling MoE gate epilogue on Intel XPU
- [Bug] [ROCm] DeepSeek-V4-Flash FP8 on MI300X (gfx942): output degrades with context length — coherent <100 tokens, garbage by ~300, NIAH 0/3
- Allow platforms to declare Mamba extra-buffer support
- Docs
- Python not yet supported