sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- fix: show bs alongside num_tokens in CUDA graph capture/compile progress
- [Fix] profiler: remove type=bool from BooleanOptionalAction args (crashes on Python 3.14)
- [Bugfix] Defer beam initialization until request admission
- [AMD][CI] Retire legacy GSM8K/MMLU from AMD CI in favor of sgl-eval
- fix(profiler): handle memory snapshot failures
- [Model] Add native block FP8 loading for K2 Horizon MoVA
- Fix compressed-tensors MoE zero-point layout conversion
- [Unified Cache]: Add hybrid mamba support to the external-cache linker
- [Bug] NEXTN with Qwen3.5-122B-A10B-FP8 memory-faults on MI300X/MI325X/MI355X (aiter backend, rocm10 nightlies); triton works, SGLANG_USE_AITER_UNIFIED_ATTN works on gfx942 only
- [Bugfix] Honor per-module quantization settings in Persimmon
- Docs
- Python not yet supported