sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Bug] Hardcoded fp32 dtypes in KDA short-convs cause 2x bandwidth waste
- [Spec Decode] Add candidate-pruned DSpark Markov sampling
- [NPU] Add optional ScatterPa KV cache writes
- [Bugfix] Respect checkpoint dtype for KDA short-conv weights
- fix(cuda-graph): remove unnecessary CP batch alignment
- [NPU] Fix MLA HiCache initialization with MTP enabled
- [Bug] sm_120 - tvm.error.InternalError: Error in function 'TllmGenFmhaRunner' at sglang/lib/python3.12/site-packages/flashinfer/data/include/flashinfer/trtllm/fmha/fmhaRunner.cuh:37: Unsupported architecture
- [Bug][Diffusion] Warmup request finalization reloads offloaded components and triggers OOM on Wan2.2 A14B
- [feat] Support DeepSeek-V4-Flash / DeepSeek-V4-Flash-0731 on SM89
- [KV Cache] Align additive KV event metadata with shared consumers
- Docs
- Python not yet supported