sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Bug] Qwen3-VL vision features diverge from Transformers/vLLM in v0.5.17 on fine-grained grounding
- [Kernel] Optimize Triton decode attention with hardware-aware dynamic Split-K scaling (#2271)
- [Playground] Verified cell: Qwen3.8-27B / dgx-spark / nvfp4 / DFLASH2 / single
- [Bug] /health handler's timeout path doesn't cancel the scheduler-side request, causing orphaned health-check requests to pile up and crash paged-prefill batching
- fix: bypass Qwen3.5 fused CUDA preparation for mRoPE
- [Bug] DSV4F0731+DSPARK ERROR IN SGLANG 0.5.18
- [CPU][Diffusion] Add packed QKV kernel and enable Ulysses SP on CPU
- [weight cache] Add disk-vs-daemon parity harness; fix persistence, zombie-PID liveness, and SM100 block-FP8 gating
- fix(constrained): require output after reasoning
- [Bug] sglang-kernel JIT: HiCache L3 segfault on first backup with nvcc 12.8 + CUDA 13 runtime (compile-time vs runtime cudaMemcpyBatchAsync ABI in staged_write_back.cuh)
- Docs
- Python not yet supported