sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Fix] Scope the GLM47 tool grammar to requests that allow tool calls
- [dLLM] Support per-request temperature, top-p and top-k sampling
- [AMD] Let the GLM DSA NextN draft declare its own shared-expert fusion architecture
- [Frontend] Parallelize long chat prompt encoding
- [MLX] Compute prefill logits for the last position only
- [AMD][ROCm] Keep cos_sin_cache fp32 on HIP for fused QSA indexer kernel
- [Fix] Feed the relayed last token to penalizers under the overlap scheduler
- [Perf] Use the fused low-ratio compressor on Hopper
- [CPU] Add FP8 quantization support for compressed-tensors paths
- [CPU] Add INT8 quantization support for per-token and compressed-tensors W8A8 paths
- Docs
- Python not yet supported