sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Test] Pin DSpark PD decode draft-input handoff regression
- Fix DSA metadata row count for DP-padded idle speculative batches
- [Bug] MI355X Qwen3.5 MTP throughput is significantly behind B200/B300 on realistic agentic workloads
- Prevent Qwen3.5 MTP draft from inheriting GPTQ quantization
- [Bug] Hybrid Mamba prefill allocation failure kills scheduler instead of returning request to waiting queue
- [Bug] /v1/responses: `created_at` is a float in streaming events but an int in non-streaming responses
- [Bug] DeepSeek-V4 sparse attention indexer (`fp8_paged_mqa_logits`) illegal memory access with long-context requests
- [Bug] Scheduler crashes with AttributeError ('list' object has no attribute 'tolist') on mixed batches with token_ids_logprob — prefill and decode paths, v0.5.14–v0.5.17
- [kernel] One rmsnorm kernel for every hidden size, tuned from Python
- Widen swapAB dispatch range in SM120 fp8 blockwise GEMM
- Docs
- Python not yet supported