sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Spec][Args] Add --speculative-draft-json-model-override-args to override the draft model config independently of the target's
- [NPU] Introduce minimal DeviceOperator for runtime layout adaptation
- feat(minimax): integrate FlashInfer MSA on Blackwell
- [DeepSeek-V4][DSpark] Handle zero-length CUDA graph padding in online C128 planner
- # Breakable CUDA graph: per-layer break dominates GDN prefill time on hybrid models (Qwen3.5/3.6)
- fix: handle list token-id logprob results
- fix(kimi-k3): align reported cached tokens with prompt usage
- bug fix
- Add NIXL HiCache Storage Metrics
- [FlashInfer] perf: Fuse FP4 MoE input quantization and shared-expert SwiGLU
- Docs
- Python not yet supported