sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [AMD] Fix RMSNorm forward_hip on vLLM's in-place fused_add_rms_norm
- [AMD] DeepSeek-V4 on gfx942: correct FP8 and routing numerics
- fix(rocm): DeepSeek-V4 asymmetric SwiGLU clamp in the HIP Triton MoE path
- fix(rocm): make in-place reloads of gfx94x block-FP8 weights and FP8 KV-cache scales work
- [sglang-miles] Restore FP8 checkpoint semantics on every in-place reload path (gfx94x)
- [sgl-router] Discover models for keyless external workers
- [Tokenizer] Qualify the multi-tokenizer args shm name with the PID namespace
- [CI] ltx_2_3_hq_pipeline perf checks fail on most diffusion PRs (load / decode / denoise variance)
- [HiCache] Store the MHA L2 host pool as INT8, 1.78x more cached tokens per GB
- [Fix] Respect an explicit SGLANG_FLASHINFER_WORKSPACE_SIZE
- Docs
- Python not yet supported