sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Weight Cache] Support daemon-to-daemon heterogeneous weight resharding with Mooncake backend
- feat(hicache): export prefetch outcome metrics
- [kernel] Optimize MiniMax sparse KV page lookups
- fix(gdn): normalize tracked conv state dtype
- --enable-deterministic-inference accepted and reported enabled, but has no effect on an NVFP4 W4A4 + hybrid-Mamba checkpoint with a DFlash2 drafter (GB10 / sm_121)
- [Daemon weight Cache] Daemon To Daemon transfer backend
- [Kimi-K3] Share replicated-Q weight storage
- [Bug] MegaMoE a2a backend crashes with CUDA illegal address on multi-node TP (SM90): SymBuffer peer access assumes a single NVLink domain
- [Kimi-K3] Enable fused KDA decode for TP4
- [Perf][DSA] Unify q8kv8 sparse-prefill topk_length source with the bf16 path; delete backscan kernel
- Docs
- Python not yet supported