Benchmarks · no Elo
SubagentBench
Best model for each sub-agent task (explore, webfetch, compaction), on real prompts from our todos.
scoring: judge scores · todo-bench/
Explore sub-agent
r5 · 2026-09-23 · 2 runs each20 real code-search questions our agents asked, picked from 48 explore calls for variety.
ModelScoreHalluc. / runTimeCost (list price, per run)
GPT-6 Sol8.930.002m 37s$0.023
Claude Opus 5.58.800.332m 20s$0.143
GPT-6 Luna6.650.551m 42s$0.001
Claude Haiku 4.53.454.501m 54s$0.04
Webfetch extraction
r1 · 2026-09-24177 real url + question pairs from our todos; every model reads the same page snapshot.
ModelScoreHalluc. / runTimeCost (list price, all 177)
GPT-6 Sol9.210.056s$0.36
Claude Opus 5.58.610.548s$2.51
GPT-6 Luna8.510.235s$0.02
Claude Haiku 4.56.811.104s$0.42
Compaction summary
r1 · 2026-09-3020 real long conversations. The question: can a fresh agent continue from the summary alone?
ModelScoreHalluc. / runTimeCost (list price)
GPT-6 Sol9.000.054m 27s—
Claude Sonnet 5.58.750.451m 20s—
GPT-6 Luna7.830.821m 46s—
Claude Sonnet 57.002.102m 1s—
Claude Haiku 4.53.955.0543s—
Scores 1–10 from two blinded judges (Claude Opus 5.5 and GPT-6 Sol); both agree on the order. Time is the median per call.