Terminal-Bench 2.1 at 95.1% with Opus 5.5: how the number was produced
77 of the 81 tasks Opus 5.5 agreed to attempt; 8 safety refusals excluded (86.5% counting them). Plus the earlier GPT-5.6 Sol runs, the rerun rule, cost accounting, and why none of it is a leaderboard entry.