183
sourced data points across two hubs
12 headline + 171 sectioned
47
primary sources
surveys, RCTs, telemetry, leaderboards
9
derived formulas
each computed from cited figures
82
glossary terms
90 records, 8 shared slugs

Three reference pages sit outside the chronological feed and are maintained rather than dated: two statistics hubs and one glossary. This post documents how they are built so a reader knows what a number on them means and how far it can be trusted.

The two hubs

HubSectionsData pointsSourcesUpdated
AI agent statisticsadoption · ROI · developers · capability · reliability · workforce92212026-09-13
AI coding tool statisticsadoption · productivity · quality · benchmarks · agents · workforce91262026-09-13

Every figure links to a numbered source entry that records the publisher, the report name, and the date. The date is the measurement or publication date of the source, not the date the hub was updated — a survey fielded in mid-2025 stays labelled 2025 however recently the page was refreshed.

What a number on a hub is, and is not

The hubs deliberately place three kinds of evidence on the same page. They do not measure the same thing.

Survey
reported behaviour
  • "84% use or plan to use AI tools"
  • population + fielding date matter
  • self-report bias
Controlled study
measured effect
  • "+19% task time in an RCT"
  • sample and task set matter
  • narrow generalisation
Benchmark
pass rate on a fixed set
  • "61.5% on SWE-bench Pro"
  • dataset version + harness matter
  • contamination risk

A survey figure and a benchmark figure that appear to conflict usually do not; they answer different questions. The hubs keep them in separate sections and label the source type so the reader can tell which is which.

The derived formulas

Each hub carries a Formulas section with quantities computed from the cited figures rather than quoted from a source. They exist so the numbers can be reasoned with. Three examples, with their inputs:

Perception gap (coding tools hub). METR’s randomised trial measured experienced developers at +19%+19\% task time with AI while they estimated −20%-20\%:

Δpercep=t^−t=(−20%)−(+19%)=−39 pts\Delta_{\text{percep}} = \hat{t} - t = (-20\%) - (+19\%) = -39\ \text{pts}

Reliability under repetition (agents hub). If a single task succeeds with probability pip_i per independent run, the probability that it succeeds on all kk runs decays geometrically; across a benchmark the metric is the average of those per-task values, not the aggregate pass rate raised to kk. τ-bench reports passk\text{pass}^k for exactly this reason: a system that passes a task about half the time is much less than half as reliable when the same task must succeed eight times in a row.

passk=1N∑i=1Npi k\text{pass}^k = \frac{1}{N}\sum_{i=1}^{N} p_i^{\,k}

Review load multiplier (coding tools hub). Faros telemetry shows high-AI teams merging 98% more pull requests, each 154% larger:

R=(1+ΔPRs)(1+Δsize)=1.98×2.54≈5.0×R = (1 + \Delta_{\text{PRs}})(1 + \Delta_{\text{size}}) = 1.98 \times 2.54 \approx 5.0\times

Each formula card names the assumption that limits it — for the review multiplier, that both effects apply to the same teams; for passk\text{pass}^k, that runs of the same task are independent.

The glossary

The glossary has 82 term pages generated from two topic sets: 41 agent terms and 49 coding-tool terms. Eight slugs appear in both sets — context window, MCP, coding agent, SWE-bench, benchmark saturation and three more — and resolve to one URL and one definition, with the page listing both topic-set memberships, rather than to two competing entries.

Each term page carries a one-sentence definition, a longer explanation, and links back to the statistics that use the term. The sentence definition is the canonical one; the statistics hubs and the comparison pages use terms in that sense.

Citing

The hubs publish under CC BY 4.0 with a BibTeX entry at the foot of each page, and expose the source table as structured data. Cite the primary source for the figure and the hub for the compilation; a hub number without its source is a secondary citation.

For a worked example of attaching methodology to a single headline figure, see the Terminal-Bench 2.1 methodology.