sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [MLX] Cap the MLX buffer cache by default
- [Example] Serve Jev's /v1/systemone typed-decision API on top of /v1/score
- perf(dsv4.1): add virtual pipeline parallelism scheduler
- fix: partial rotary embeddings on XPU, UnifiedRadixCache default, Transformers-backend remote-code fallback, tokenizer pad() compat shim, and added-token id remap
- [Bugfix] Support LogitsProcessorOutput in the breakable CUDA graph output buffers
- [function_call] Keep text after dots tool calls in non-streaming parsing
- [ROCm] Rotate even-batch decode attention grid across XCDs
- [dLLM] Support DiffusionGemma decision API and output logprobs
- [Fix] Scope the GLM47 tool grammar to requests that allow tool calls
- [AMD] aiter unified attention: refresh the CUDA-graph page table at page_size 1
- Docs
- Python not yet supported