aobench
https://github.com/mskazemi/aobench
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
not yet supported1 Subscribers
Add a CodeTriage badge to aobench
Help out
- Issues
- fix: prevent new runs of blocked tasks and quarantine unsupported PERF cases
- Generate and send a real RBAC-hard-fail fixture to the EvalPort spec (discussion #51)
- Flaky release gate: test_report_json_flag_emits_clean_json_on_stdout fails ~1 run in 8
- Ten tasks grant tool families that do not exist — three dev tasks get zero tools, and validate passes
- Two dev-split tasks are unpassable: `my own job` is owned by someone else, so a correct agent is RBAC hard-failed
- Verify every command in the docs actually does what the page says — from a clean checkout
- [Result] Claude Sonnet Benchmarking
- Write a second task for a thin QCAT × role cell (32 of 50 have only one)
- Run a model and submit a leaderboard entry — negative results welcome
- LiteLLM adapter — one file, ~100 providers
- Docs
- not yet supported