sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- fix(grpc): release GIL during native server shutdown
- [Perf] Lossless final-layer prefill optimization for LLaMA models
- [Bugfix][VLM] Prevent server SIGQUIT by zero-padding short multimodal embeddings
- [Bug] HiCache host-memory sizing guard charges co-located ranks twice and rejects pools that fit (host_memory_budget_bytes, #35540)
- [Bug][Diffusion] MiniMax-H3: 1344x768 fails deterministically with "CUDA driver error: device not ready" on 12GB, and the failed request poisons the server
- [Quant] Select FlashInfer b12x NVFP4 GEMM by default on SM120
- [Perf] Keep torchvision, openai SDK and FastAPI off the scheduler imp…
- [Fix] set_mla_kv_buffer: reshape non-contiguous MLA latent slices
- fix(grpc): propagate cancellation and accept IPv6 bind addresses
- [gdn] Add cudnn frontend backend for GDN
- Docs
- Python not yet supported