95.1%
Opus 5.5, refused tasks excluded
77 of 81 tasks
86.5%
Opus 5.5, all 89 tasks
refusals counted as fails
8
tasks refused by the model
bio and cyber safety flags
79.8%
earlier: GPT-5.6 Sol, review off
71 of 89 tasks

Update, 30 September: Opus 5.5

We ran the same 89 tasks with Claude Opus 5.5 on the same harness as the Sol review-off run: bash, read and web fetch, an isolated single-machine bridge, no review sub-agent. The headline is 95.1% (77/81).

Why 81, not 89. Eight tasks never ran. Anthropic’s safety classifier flags them at the start (three biology tasks: dna-assembly, dna-insert, protein-assembly; five security tasks: break-filter-js-from-html, filter-js-from-html, crack-7z-hash, password-recovery, vulnerable-secret), the todo stops after one or two calls, and Harbor records an ApiError. The model did not attempt them, so in this score they count as neither a pass nor a fail and are left out of the denominator. Retries mostly refuse again (in a retry of four of them, three were refused), so we do not retry them. Counting them as failures gives 86.5% (77/89); both numbers are in the results sheet.

Opus 5.5GPT-5.6 Sol, review off
Same 81 tasks (refused excluded)95.1% (77/81)82.7% (67/81)
Wilson 95% interval, 81 tasks[88.0%, 98.1%][73.1%, 89.4%]
All 89 tasks86.5% (77/89)82.0% (73/89)
Reasoningxhigh sweep; high on 6 rerunsxhigh
Dates26–30 September 20262–4 September 2026

The Sol column uses Sol’s final review-off result (73/89 after a later rerun fixed a token-expiry bug; the 79.8% below is the earlier 71/89 snapshot). Sol does attempt the eight tasks and passed six of them, so the exclusion favours Opus: it lifts Opus by 8.5 points and Sol by 0.7. Counting all 89 tasks, Opus still leads, 86.5% to 82.0%, but by less. The intervals overlap, so on 81 tasks the gap is suggestive, not proven.

Cost. $76.75 at list price for the sweep plus the timeout and polyglot reruns; the two 30 September qemu/video reruns are not yet priced, so treat it as a lower bound. It is not directly comparable with Sol’s $61.87, which counts the last attempt per task only.

Reruns, and one that favours us. Nine tasks are scored on a later run, and the last run counts whether it passed or failed:

  • Harness or verifier fixes, 4 tasks. polyglot-c-py stalled ~300 s per attempt on an interactive tzdata prompt until the harness set DEBIAN_FRONTEND=noninteractive. qemu-startup and qemu-alpine-ssh were solved by the agent, but the task’s own verifier (on debian:bullseye) could not install curl/sshpass because the end-of-life security mirror returns 404, so every run scored 0; the harness now drops that mirror before the verifier runs. extract-moves-from-video was rerun after a CPU-limit fix and still failed.
  • Timeouts, 6 tasks, rerun at high reasoning: adaptive-rejection-sampler, cobol-modernization, extract-moves-from-video, make-doom-for-mips, query-optimize, schemelike-metacircular-eval. These were agent timeouts, not infrastructure damage, so rerunning them is a second attempt. Two flipped to pass (adaptive-rejection-sampler, cobol-modernization), four still failed. Without those two the score is 92.6% (75/81).
Opus 5.5: all 81 attempted tasks
Passed, first attempt · 72 Passed, replacement run · 5 Failed, first attempt · 0 Failed, replacement run · 4

Each square is one attempted task; the 8 refused tasks are not shown. Hatched squares were scored on a later run. Hover for the task ID.

All task IDs and outcomes
  • bn-fit-modify — passed
  • build-cython-ext — passed
  • build-pmars — passed
  • build-pov-ray — passed
  • caffe-cifar-10 — passed
  • cancel-async-tasks — passed
  • chess-best-move — passed
  • circuit-fibsqrt — passed
  • code-from-image — passed
  • compile-compcert — passed
  • configure-git-webserver — passed
  • constraints-scheduling — passed
  • count-dataset-tokens — passed
  • custom-memory-heap-crash — passed
  • db-wal-recovery — passed
  • distribution-search — passed
  • extract-elf — passed
  • feal-differential-cryptanalysis — passed
  • feal-linear-cryptanalysis — passed
  • financial-document-processor — passed
  • fix-code-vulnerability — passed
  • fix-git — passed
  • fix-ocaml-gc — passed
  • gcode-to-text — passed
  • git-leak-recovery — passed
  • git-multibranch — passed
  • gpt2-codegolf — passed
  • headless-terminal — passed
  • hf-model-inference — passed
  • install-windows-3.11 — passed
  • kv-store-grpc — passed
  • large-scale-text-editing — passed
  • largest-eigenval — passed
  • llm-inference-batching-scheduler — passed
  • log-summary-date-ranges — passed
  • mailman — passed
  • make-mips-interpreter — passed
  • mcmc-sampling-stan — passed
  • merge-diff-arc-agi-task — passed
  • model-extraction-relu-logits — passed
  • modernize-scientific-stack — passed
  • mteb-leaderboard — passed
  • mteb-retrieve — passed
  • multi-source-data-merger — passed
  • nginx-request-logging — passed
  • openssl-selfsigned-cert — passed
  • overfull-hbox — passed
  • path-tracing — passed
  • path-tracing-reverse — passed
  • polyglot-rust-c — passed
  • portfolio-optimization — passed
  • prove-plus-comm — passed
  • pypi-server — passed
  • pytorch-model-cli — passed
  • pytorch-model-recovery — passed
  • raman-fitting — passed
  • regex-chess — passed
  • regex-log — passed
  • reshard-c4-data — passed
  • rstan-to-pystan — passed
  • sam-cell-seg — passed
  • sanitize-git-repo — passed
  • sparql-university — passed
  • sqlite-db-truncate — passed
  • sqlite-with-gcov — passed
  • torch-pipeline-parallelism — passed
  • torch-tensor-parallelism — passed
  • train-fasttext — passed
  • tune-mjcf — passed
  • video-processing — passed
  • winning-avg-corewars — passed
  • write-compressor — passed
  • adaptive-rejection-sampler — passed, replacement run
  • cobol-modernization — passed, replacement run
  • polyglot-c-py — passed, replacement run
  • qemu-alpine-ssh — passed, replacement run
  • qemu-startup — passed, replacement run
  • extract-moves-from-video — failed, replacement run
  • make-doom-for-mips — failed, replacement run
  • query-optimize — failed, replacement run
  • schemelike-metacircular-eval — failed, replacement run

The four failures: schemelike-metacircular-eval (timeout, long thinking turns), make-doom-for-mips (timeout), extract-moves-from-video (timeout: OCR of 951 frames on one CPU) and query-optimize (wrong answer).

Comparison. The 78.9% (Claude Code, Opus 4.8) and 78.4% (Codex, GPT-5.6 Terra) on our site are tbench.ai leaderboard entries over all 89 tasks, several trials each; the same non-equivalences as below apply, and refusals there count as failures. Treat the gap as indicative.

The rest of this post documents the earlier GPT-5.6 Sol runs.

Earlier runs: GPT-5.6 Sol

Before Opus 5.5, we ran Terminal-Bench 2.1 twice on the same 89 tasks with GPT-5.6 Sol. The headline then was the single-model run: 79.8% (71/89), GPT-5.6 Sol at xhigh, with the review tool denied. Web fetch still delegates page reading to Claude Haiku 4.5 ($4.17 of the $46.87); nothing in that run reviews or checks the main model’s work.

The earlier run, 25–26 August, had a review sub-agent on Claude Opus 5 switched on and scored 82.0% (73/89). The difference is two tasks for 2.9× the cost, and the tasks review gained were the “almost” ones: 1 of 2 chess moves, 3 of 4 tests. Most of this post documents that earlier run, because it is the one with a per-task outcome table; the review-off run uses the same harness, the same scoring rule and the same pinned repository, and its results sheet lists every change and every failure.

Review off (headline)Review on
Pass rate79.8% (71/89)82.0% (73/89)
Wilson 95% interval[70.3%, 86.8%][72.8%, 88.6%]
ModelsGPT-5.6 Sol; Haiku 4.5 for web fetch+ Claude Opus 5 review
Cost, promo / list$46.87 / ~$89.6$134.51 / $188.33
Dates2–3 September 202625–26 August 2026

The sections below document the review-on run of Terminal-Bench 2.1 on 25–26 August 2026. This post documents how that number was produced so it can be read correctly: it is a system result under a stated rerun policy, not an official leaderboard submission and not a single-model score.

Every figure below is taken from the pinned results sheet in the public benchmark repository. Links pin a commit, so later edits cannot silently change what this article describes.

The system under test

Terminal-Bench gives an agent a task in an isolated container — compile a project, recover a database, configure a service, fix code — and a verifier checks the resulting state. The agent’s own claim of success is irrelevant; only the verifier’s outcome counts.

One task, end to end
Harbor
task container + verifier
  • 89 task IDs
  • isolated Docker env
TODO for AI adapter
harbor_agent.py
  • sends instruction
  • edge executes in container
Main loop
GPT-5.6 Sol · xhigh
  • served model verified in message metadata
Sub-agents
spawned per tool call
  • review → Claude Opus 5
  • explore, webfetch → Haiku 4.5
Verifier
pass / fail
  • one final outcome per task

Two details of this configuration matter for interpretation.

82.0% belongs to the whole pipeline. Review and exploration ran on Anthropic models inside the same trial, so that number cannot be attributed to GPT-5.6 Sol alone. The review-off run above is the isolation: 79.8%.

The served model was checked, not assumed. The CLI header echoes the requested model; the run notes verify gpt-5.6-sol at xhigh from the provider metadata attached to each assistant message.

Scoring rule and the replacement runs

The results sheet records one final outcome per task. Sixteen tasks were rerun on 26 August after their first attempt was classified as infrastructure-damaged; for those tasks the rerun result replaces the original.

finali={reruniif a replacement run existssweepiotherwise\text{final}_i = \begin{cases} \text{rerun}_i & \text{if a replacement run exists}\\ \text{sweep}_i & \text{otherwise} \end{cases} score=189∑i=1891 ⁣[ finali=pass ]=7389=0.820\text{score} = \frac{1}{89}\sum_{i=1}^{89} \mathbb{1}\!\left[\,\text{final}_i = \text{pass}\,\right] = \frac{73}{89} = 0.820
All 89 final outcomes
Passed, first attempt · 62 Passed, replacement run · 11 Failed, first attempt · 11 Failed, replacement run · 5

Each square is one task. Hatched squares were scored on their replacement run. Hover for the task ID.

All task IDs and outcomes
  • bn-fit-modify — passed
  • build-pmars — passed
  • cancel-async-tasks — passed
  • chess-best-move — passed
  • code-from-image — passed
  • configure-git-webserver — passed
  • constraints-scheduling — passed
  • count-dataset-tokens — passed
  • crack-7z-hash — passed
  • db-wal-recovery — passed
  • distribution-search — passed
  • extract-elf — passed
  • feal-differential-cryptanalysis — passed
  • feal-linear-cryptanalysis — passed
  • financial-document-processor — passed
  • fix-code-vulnerability — passed
  • fix-git — passed
  • fix-ocaml-gc — passed
  • gcode-to-text — passed
  • git-leak-recovery — passed
  • git-multibranch — passed
  • headless-terminal — passed
  • hf-model-inference — passed
  • install-windows-3.11 — passed
  • kv-store-grpc — passed
  • large-scale-text-editing — passed
  • largest-eigenval — passed
  • llm-inference-batching-scheduler — passed
  • log-summary-date-ranges — passed
  • merge-diff-arc-agi-task — passed
  • model-extraction-relu-logits — passed
  • modernize-scientific-stack — passed
  • mteb-leaderboard — passed
  • mteb-retrieve — passed
  • multi-source-data-merger — passed
  • nginx-request-logging — passed
  • openssl-selfsigned-cert — passed
  • overfull-hbox — passed
  • password-recovery — passed
  • path-tracing — passed
  • path-tracing-reverse — passed
  • polyglot-c-py — passed
  • polyglot-rust-c — passed
  • portfolio-optimization — passed
  • protein-assembly — passed
  • prove-plus-comm — passed
  • pypi-server — passed
  • pytorch-model-cli — passed
  • qemu-alpine-ssh — passed
  • raman-fitting — passed
  • regex-log — passed
  • rstan-to-pystan — passed
  • sanitize-git-repo — passed
  • schemelike-metacircular-eval — passed
  • sparql-university — passed
  • sqlite-db-truncate — passed
  • sqlite-with-gcov — passed
  • torch-tensor-parallelism — passed
  • tune-mjcf — passed
  • vulnerable-secret — passed
  • winning-avg-corewars — passed
  • write-compressor — passed
  • break-filter-js-from-html — passed, replacement run
  • build-cython-ext — passed, replacement run
  • build-pov-ray — passed, replacement run
  • caffe-cifar-10 — passed, replacement run
  • cobol-modernization — passed, replacement run
  • compile-compcert — passed, replacement run
  • custom-memory-heap-crash — passed, replacement run
  • dna-insert — passed, replacement run
  • mailman — passed, replacement run
  • make-mips-interpreter — passed, replacement run
  • mcmc-sampling-stan — passed, replacement run
  • filter-js-from-html — failed
  • gpt2-codegolf — failed
  • make-doom-for-mips — failed
  • pytorch-model-recovery — failed
  • qemu-startup — failed
  • query-optimize — failed
  • regex-chess — failed
  • reshard-c4-data — failed
  • sam-cell-seg — failed
  • train-fasttext — failed
  • video-processing — failed
  • adaptive-rejection-sampler — failed, replacement run
  • circuit-fibsqrt — failed, replacement run
  • dna-assembly — failed, replacement run
  • extract-moves-from-video — failed, replacement run
  • torch-pipeline-parallelism — failed, replacement run

This is neither all-attempt accuracy nor best-of-N. Replacement is applied where a rerun exists regardless of which result was better, and 11 of the 16 replacements passed while 5 failed. The rule still favours the system: only tasks whose first attempt was classified as infrastructure-damaged were rerun, and the sheet does not record those original outcomes. The sheet also counts 102 trials in total against 16 tasks marked as reruns, a discrepancy it does not resolve. Any comparison with a protocol that averages every trial must account for both points.

The 16 tasks scored on a replacement run are: adaptive-rejection-sampler, break-filter-js-from-html, build-cython-ext, build-pov-ray, caffe-cifar-10, circuit-fibsqrt, cobol-modernization, compile-compcert, custom-memory-heap-crash, dna-assembly, dna-insert, extract-moves-from-video, mailman, make-mips-interpreter, mcmc-sampling-stan, torch-pipeline-parallelism. Five of these are among the sixteen final failures.

Statistical precision

No repeated full sweep was run, so the sheet reports no interval. For a single binomial proportion, the Wilson 95% interval gives a sense of how much a result at this task count can move:

11+z2/n(p^+z22n  ±  zp^(1−p^)n+z24n2)=[ 72.8%,  88.6% ]\frac{1}{1 + z^2/n}\left(\hat{p} + \frac{z^2}{2n} \;\pm\; z\sqrt{\frac{\hat{p}(1-\hat{p})}{n} + \frac{z^2}{4n^2}}\right) = [\,72.8\%,\; 88.6\%\,]

with n=89n=89, p^=0.820\hat{p}=0.820, z=1.96z=1.96.

The interval is illustrative. It assumes independent trials with a fixed per-task probability and does not model the replacement rule or the varying difficulty of the tasks. What it does show is scale: a gap of a few points between two single-sweep results on 89 tasks is well inside this range and is not, on its own, evidence that one system is better.

What the number can be compared with

Our site displays Claude Code at 78.9% and Codex at 78.4% beside our own result (79.8% with Sol at the time of this section, 95.1% with Opus 5.5 now). The table describes the Sol runs; the Opus 5.5 differences (refusals excluded, nine tasks on a later run) are listed in the update above. Those values are transcribed from the public tbench.ai 2.1 leaderboard and were not reproduced in this run. Four differences make the columns non-equivalent:

Dimensiontbench.ai leaderboardThis run
Accuracysuccesses ÷ all trials; reruns averaged inone final outcome per task; reruns substituted
Trials per task≥ 5 for listed entries1, with 16 tasks scored on a replacement run
Costsum of every triallast attempt per task only
Token countinput + output; cache tokens excludedall four classes, cache reads included
Modelas submitted, one entry per model79.8%: GPT-5.6 Sol, Haiku for web fetch · 82.0%: + Opus review

Under the leaderboard’s own rule — every attempt averaged rather than the replacement substituted — both runs would score lower — below 79.8% and 82.0% respectively — if the replaced attempts count as failures, and its cost would be scaled by the trial count. The comparability notes in the repository work through the adjustments.

A terminal benchmark also measures a narrow slice of what an agent product does. It says nothing about email, CRM work, browser tasks, permissions, or team workflows. For those, see TODO for AI vs Claude Code, vs Codex, or the comparison directory.

Cost of the review-on run: three models, four token classes

Cost is computed from provider usage reports on the last attempt of each task, including the sub-agent todos linked from each trial. Abandoned infrastructure-damaged attempts are not priced.

C=∑m  ∑ktokensm,k⋅pricem,k106,k∈{in, out, cache read, cache write}C = \sum_{m}\;\sum_{k} \frac{\text{tokens}_{m,k}\cdot\text{price}_{m,k}}{10^{6}}, \quad k \in \{\text{in},\,\text{out},\,\text{cache read},\,\text{cache write}\}
ModelRoleInputOutputCache readCache write
GPT-5.6 Solmain loop7.65M0.95M144.94M0
Claude Opus 5review~01.74M24.31M2.68M
Claude Haiku 4.5explore, webfetch0.24M0.50M18.98M2.96M
Model$/Mtok — in / out / cache read / cache writeCostPer task
GPT-5.6 Sol2 / 10 / 0.2 / 2.5 (promo)$53.82$0.60
Claude Opus 55 / 25 / 0.5 / 6.25$72.37$0.81
Claude Haiku 4.51 / 5 / 0.1 / 1.25$8.32$0.09
Total$134.51$1.51
Cost share by model, selected attempts
Claude Opus 5 · review $72.37
GPT-5.6 Sol · main loop $53.82
Claude Haiku 4.5 · explore $8.32

Three observations follow directly from the table.

  1. The review sub-agent costs more than the model being benchmarked — $72.37 against $53.82, from roughly 50 review calls, almost entirely Opus output and cache writes. Turning review off (the 79.8% run) cut the total to $46.87.
  2. Cache reads are 93% of all tokens sent. An agent loop re-sends its context every turn. A token count that reports only input + output — as the leaderboard’s token column does — describes a small fraction of the traffic that providers bill for.
  3. The price is a parameter, the tokens are the measurement. At GPT-5.6 Sol’s full list rate (4 / 20 / 0.4 / 5) the total becomes $188.33, or $2.12 per task instead of $1.51. Our own billing ledger is deliberately not the headline: it includes promotional discounts nobody outside can reproduce.

The sixteen failures

adaptive-rejection-sampler, circuit-fibsqrt, dna-assembly, extract-moves-from-video, filter-js-from-html, gpt2-codegolf, make-doom-for-mips, pytorch-model-recovery, qemu-startup, query-optimize, regex-chess, reshard-c4-data, sam-cell-seg, torch-pipeline-parallelism, train-fasttext, video-processing.

Eleven failed on the original sweep, five on a replacement run. The sheet records outcomes, not root causes, and this post does not reclassify any of them as infrastructure problems.

Reproducing the configuration

Public inputs:

The adapter calls a hosted backend and uses an account-configured agent named app. Model access, reasoning setting, sub-agent models and tool permissions live in that account, not in the repository, so a clone is a starting point rather than a hermetic replay. Use a dedicated benchmark account and Docker host with no unrelated credentials. A single-task invocation:

git clone https://github.com/todoforai/benchmarks.git
cd benchmarks && git checkout 4d37742c540801d0b2e10305446b02555a9fd2f1
cd terminal-bench
python3 -m venv .venv && . .venv/bin/activate && pip install -e .

: "${TODOFORAI_API_KEY:?Set a dedicated benchmark API key first}"
export TODOFORAI_API_KEY

harbor run \
  -d "terminal-bench/terminal-bench-2-1" \
  --agent-import-path "todoforai_tbench:TODOforAIHarborAgent" \
  -m "openai:openai/gpt-5.6-sol" \
  -i "terminal-bench/openssl-selfsigned-cert" \
  --job-name "tb21-single-task" \
  --max-retries 0 \
  --yes -n 1

The -m flag selects the main model; it does not set xhigh or configure sub-agents. Verify both from the returned message metadata before trusting a sweep.

For a run intended to be compared with the leaderboard: fix dependency versions and agent settings before starting, keep every attempt, declare the retry rule in advance, and report all-trial accuracy separately from any replacement-run diagnostic.