sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Feature] Token-aware admission cap (--max-inflight-prefill-tokens)
- [Bug] Cancelling the tail request in a chunked-prefill batch still leaks one visible token and leaves negative scheduler output ids
- [Diffusion] Enable vectorized JointThreshold decoding on CUDA
- [Bug] 1M-token prefill kills the engine with CUDA OOM in DSV4 indexer fp8_mqa_logits (nonpaged path) under --tp 8 + MegaMoE on 8x B200 (v0.5.17); equivalent request serves under tp8/dp8 dp-attention
- [BugFix] LoRA registry leak on request abort
- [Bug] Kimi-K3 - cross prompt reasoning leakage
- [BugFix] Synchronize Breakable CUDA Graph after warmup before capture
- fix(security): harden SafeUnpickler with exact-name allowlist for generic modules
- [Bug] DSpark compact ragged CUDA Graph uses incompatible request-slot geometry for the same token tier
- Remove many scheduler synchronization to enable overlap scheduler on agent workload
- Docs
- Python not yet supported