Benchmarks · our own harness · no Elo, no votes

TerminalBench

Terminal-Bench 2.1 tasks run through our own harness (Harbor adapter), one sweep.

scoring: pass rate · terminal-bench/

RunPassCost promo / list
GPT-5.6 Sol (xhigh), review off71/89 = 79.8%$46.87 / $89.60
same + Claude Opus 5 review sub-agent73/89 = 82.0%$134.51 / $188.33

Own harness, one sweep, infrastructure-damaged tasks replaced by their rerun. Not a tbench.ai leaderboard submission and not scored by its rules. Methodology.

Each model runs in its own sandbox through the TODO for AI agent — identical prompt, no view of the others. We publish the raw outputs, what they cost and how long they took. Nothing is voted on, nothing is fitted to a curve. Reproducible from the benchmarks repo.