sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- fix(sm120): fall back to Triton per call when FlashInfer lacks the prefill shape
- docs: clarify FLUX.2 Klein FP8 compatibility
- [Bugfix] Isolate auto-truncated token budgets in batched requests
- [Bugfix] Fix batched cross-encoder tokenization for reranking
- [Bug] GDN packed decode and KDA replay decode round sigmoid(beta) through bf16; recurrent-state drift accumulates with context length
- [FLA] Keep sigmoid(beta) in fp32 in GDN packed decode and KDA replay decode kernels
- [Bug] Kimi-K3: aborting a request mid-prefill kills every rank — `free_kv_row_segments` asserts "segment at N shares a page with the one ending at N"
- [mem_cache] Coalesce touching kv-row segments before free so an aborted mid-prefill request does not trip _page_disjoint
- [UT][NPU] Add unit tests for MoE
- [Refactor] Share `ChatInputProcessor` across API handlers
- Docs
- Python not yet supported