sglang
https://github.com/sgl-project/sglang
Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported1 Subscribers
Add a CodeTriage badge to sglang
Help out
- Issues
- [Bug] fa3 backend slow with mla page-size 64 for H20
- feat(tokenizer): segment encode_batch for long RAG (w_cap=len/1k*2, prefill cache-hit @ L20)
- [Dashboard] Add token-weighted prefix cache hit rate panel
- fix(hicache): TP/PP write-through & load-back consensus (#28429)
- [HiCache] Support configurable prefetch IO workers
- [Feature] RFC: SGLang KV Indexer for Distributed KV Cache Placement Metadata
- fix(parser): CohereCommand4Detector.force_nonempty_content emits a truncated fragment
- [Spec] perf: Fuse topk=1 target verify finalization
- Fix shared-expert TP1 double-count on skipped post-experts all-reduce
- [Bug] eagle: broadcast finalized verify decision across TP ranks (#31071)
- Docs
- Python not yet supported