sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Kernel] Selectively dequantize FP8 KV for BF16 DSA prefill
- [Bugfix] Select JIT cudaMemcpyBatchAsync ABI at runtime
- [Spec] Support PP x PD x DSpark for DSv4
- [Perf] Vectorize logit_bias construction in SamplingBatchInfo
- [Fix][DeepSeek V4] swa chunked prefill deadlock
- [Feature] Support expert-wise FP8/W4A16_AWQ mixed precision for ModelOpt MoE checkpoints
- [Fix] Clamp tombstoned v2p entries in the fused decode-metadata kernel
- [Bug] SGLANG_RAGGED_VERIFY_MODE=compact crashes decode CUDA-graph capture on Qwen3.5 hybrid linear attention: IMA in extend_attention_fwd (triton_backend)
- Support standard Content-Encoding for request decompression
- [kernel] DSA fused quant+store: per-block fp8 quantize k_nope + bf16 k_rope + paged store (lossless)
- Docs
- Python not yet supported