Benchmarks · our own harness · no Elo, no votes

TODO Bench

Our own agent pieces (explore, webfetch, compaction) on real prompts from our todos.

scoring: judge scores · todo-bench/

First runs are in progress — results land here as soon as a full sweep is done.
Each model runs in its own sandbox through the TODO for AI agent — identical prompt, no view of the others. We publish the raw outputs, what they cost and how long they took. Nothing is voted on, nothing is fitted to a curve. Reproducible from the benchmarks repo.