sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- opt: bind native Qwen MTP embed/lm_head before memory pool allocation
- feat: add Mooncake GPU-direct SpecForge capture
- [Spec] Fail loudly when the DFlash2 selector top_k exceeds the org vocab
- [Spec] Resolve DFlash2 selector shard metadata once for both TP modes
- fix(jit): pin nvcc to the fingerprinted host compiler
- Return 400 for multimodal token/data mismatches
- fix(function_call): preserve leading JSON in llama3 parser when no tool call is found
- Fix stale batch fullness after chunked prefill
- fix: select FlashInfer CUTLASS for ModelOpt mixed MoE on SM120
- [RFC] Integrating Ascend MemCache as an HiCache L3 Backend
- Docs
- Python not yet supported