Draft

Primary Sources: AI Contributions to Cyber, Math, and Algorithms

Author
Affiliation

Tom Cunningham

METR

Published

July 23, 2026

` markers; edit the entry and re-run the script. –>

This is the source log for AI’s Contributions to Cyber, Math, and Algorithms. It gathers and summarizes the primary sources in one place so that the argument document can stay short and a later reader can re-check and update each source independently of the argument.

It has grown past the point where it serves only that argument. What it is now is a reference for a broader question — how much is AI actually contributing to research and development, and how would anyone know — assembled at the level of individual checkable figures. Read as a quarry, not as a summary: it deliberately holds more material than the argument uses, records figures that turned out to be wrong, and quotes at length rather than paraphrasing.

What this document is for

The split between this file and the argument exists so that facts and inferences decay separately. An argument can be rewritten without re-checking its evidence, and a figure can be corrected without rewriting the argument. Three desiderata follow from that, in priority order.

A later reader must be able to re-check any figure without re-doing the search. Every entry carries links to primary sources, the date the figure refers to, and the date it was checked. If a number cannot be traced to something a reader can open, it does not belong in the argument.

Uncertainty must be visible at the point of use, not averaged away. Each entry states how far it has been verified and who produced it. A vendor self-report and an independent replication are both admissible; conflating them is not. Where two versions of a source disagree, both figures are recorded and the disagreement is left open rather than silently resolved.

The log must answer the questions the argument actually asks. The argument turns on eight questions, so each entry names which of them it bears on. An entry that bears on none of them is either miscategorized or should not be here.

The eight questions

Entries are tagged with the questions they speak to, using these short keys. The full versions, with each theory’s predicted answer, are in the argument’s “Four theories” table. The third column is this log’s own statement of what a satisfying answer would have to look like, and it is here because most entries fall short of it in a way worth naming: the gap between what a source measures and what the question asks is usually larger than the gap between competing sources.

Key Question What would answer it
Q1 growth rate Does AI raise the rate of efficiency growth, and by how much? A dated efficiency series for one domain, in fixed units, spanning the arrival of AI contributions, with the pre-AI slope estimated from the same series
Q2 autonomy Can AI produce a complete result without a human driving it? A result with the human contribution itemized: who chose the target, who built the scaffold, who verified, and what was left for the model
Q3 demand What happens to demand for human researchers? Employment or wages for a research occupation, with a counterfactual — not stated hiring intentions, not job-posting language
Q4 expertise How expert are the people who make discoveries with AI? Discovery outcomes by user expertise, randomized, in a research setting rather than a task-execution setting
Q5 returns How fast do returns to AI spend diminish against human labor? Repeated equal-budget runs against a fixed target population, reporting the yield of each successive run
Q6 intertemporal Is it better to spend on AI now or later? Whether last period’s spend lowered this period’s yield on the same targets, with human effort on those targets as the control
Q7 incidence Which fields and starting points does AI affect most? One codebase or problem set staged at two levels of prior optimization, with model, scaffold, budget, and metric held fixed
Q8 benchmarks Which benchmarks predict real value? A benchmark score and a realized deployment outcome for the same system, so the correlation can be estimated rather than asserted

How to use this log

To check a figure the argument uses, follow its #src- anchor and read the entry’s quotes and status line. If the status is unverified or vendor, the argument should not be resting weight on it, and the entry says what a reader would have to do to fix that.

To answer one of the eight questions from scratch, use the question map below rather than the reading order. Every entry declares which questions it bears on, so the map collects them; the counts are uneven in a way that is itself a finding, and the “Known gaps” section at the foot says what the thin rows are missing.

To assess a claim about AI’s contribution to R&D that came from somewhere else, three cross-cutting syntheses assembled here are the fastest route: the inventory of every rate that has a denominator, the inventory of every cost-per-result figure, and the coded ladder of what “autonomous” turned out to mean in each case where it was claimed. Those three catch most of what goes wrong with a headline.

To extend it, the standing rules below are the format, and the validator enforces the parts of it that can be checked mechanically. The rules that cannot be checked mechanically — quote the source rather than summarizing it, prefer a denominator to a headline, keep withdrawn figures rather than deleting them — are the ones that matter most.

Weighing the evidence

Entries are not interchangeable, and the log’s status vocabulary tracks only whether a figure was read correctly, not whether the study design can support the claim. Those are different questions and the second one does more damage when it is skipped. This is the log’s own ranking of designs, roughly from strongest to weakest for the purpose of estimating AI’s contribution to research.

Randomized assignment of AI access, with a real task and a measured outcome. Only a handful of entries here qualify, and they disagree in sign, which is the most useful thing about them. A randomized trial still cannot tell you about a setting it did not sample, and every one of these samples task execution rather than discovery.

Staged or phased deployment with a plausible counterfactual. Weaker than randomization but usually larger and more realistic. The identification rests on the rollout being unrelated to the outcome, which is an assumption rather than a fact.

A fixed benchmark measured at two dates with the structure held constant. This is what makes a capability series interpretable. Almost nothing here manages it, because scaffold, prompt, budget, and task set move together with model generation; where an entry does manage it, the entry says so, and those are the load-bearing capability comparisons in the log.

A benchmark score with a denominator someone outside the vendor can reproduce. Cost per task, human-hours to complete, success out of a stated problem count. The denominator is what makes a score comparable to anything else, and its absence is the single most common defect in this literature.

An operational log or leaderboard. Submission counts, duplicate rates, records with dated authorship. Not designed to answer anything, which cuts both ways: no researcher chose the outcome measure to flatter a conclusion, and no researcher controlled the confounds either.

A demonstration, on a target the demonstrator chose. Nearly every headline result in this log is one of these. A demonstration establishes that something is possible and says nothing about a rate, because the denominator — how many targets were tried, at what cost, with what failures — is normally unpublished. Vendor demonstrations additionally select on outcome before publication.

A survey of stated perceptions or intentions. Included where nothing better exists, and flagged. The one entry here that measured perception against measurement in the same population found the perception wrong in sign [→ METR RCT], which is the reason to discount this row heavily rather than merely to caveat it.

Two corollaries the log tries to hold to. A figure’s provenance and its verification status are orthogonal: a vendor claim can be verified as having been made while remaining unaudited as a measurement, and the status line keeps those apart. And an absence of evidence in a thin row is not evidence of a null — the “Known gaps” section exists so that silence does not read as a finding.

What an entry must contain

  • A level-3 heading ending with the year in parentheses, and a stable {#src-...} anchor. The year is the year the work first became public, not the year of a later journal version or of retrieval, and a range for an ongoing series. Anchors are cited from the argument and from other entries, so renaming one means updating both in the same commit.
  • A leading italic line giving the source type: independent, government, or vendor.
  • Key figures as bulleted claims, each dated where the date matters, and each carrying the source’s own words. See the rule on quotation below — it is the most demanding requirement here and the one most often skipped.
  • A **Dates:** line. Every entry has one; see the rule on dates below.
  • A **Bears on:** line naming the applicable questions.
  • A Links: line reaching primary sources wherever they exist, with skeptical or contradicting coverage linked alongside rather than omitted. A derived entry has no external source, so it names its input entries by anchor instead.
  • A **Status:** line, using the vocabulary below. It must say how far the figures have been checked, not only who produced them — provenance and verification are different claims, and an entry giving only the first cannot be classified.
  • A **Unused:** line, if the argument does not yet cite the entry. See the rule on holding material below.
  • A figure, where one exists that carries the entry’s point better than the prose. See the rule on figures below.

Six of these are checked mechanically: the anchor’s uniqueness, the year in the heading, the Dates line and that it names a year, the Status line and its vocabulary, the presence of links or of named inputs, and that every #src- cross-reference resolves in both directions between this file and the argument. The generated tables above are also checked for staleness. Run python3 tools/qmd_validate.py --qmd posts/2026-07-23-apple-picking-cyber-math-optimization.llm.qmd. Nothing checks the quotations, the denominators, or the dates themselves, which is where the real work is.

Status vocabulary

  • verified — checked against the primary page or paper.
  • verified-abstract — checked against an abstract, summary page, or leaderboard only; body figures not read in full.
  • unverified — surfaced by search and not yet checked against a primary source.
  • vendor — self-report. Orthogonal to the other three: a vendor claim can be verified as having been made while remaining unaudited as a measurement.
  • derived — not a source at all, but a comparison or computation this log assembled from entries that are. A derived entry exists so the argument can cite a synthesis once instead of rebuilding it, and it must name every input entry and be reproducible from them. Nothing in a derived entry may be quoted as though a source said it.

Standing rules

Prefer a denominator to a headline. “Found N bugs” is nearly uninterpretable; “N confirmed from M submissions” or “N at $X per attempt” can be compared across sources. Where only a numerator is available, say so.

Record withdrawn and dropped figures rather than deleting them. A figure that was removed from a draft, or a paper that has been disavowed, is information a later reader needs — otherwise the next person re-finds it and re-uses it. Keep the entry and mark it.

Do not let a benchmark’s own framing set the entry’s framing. Several sources here disclaim being benchmarks, revise their own problem sets, or rank differently under different scoring rules. Those caveats belong in the entry, not in a footnote of the argument.

Quote the source, do not summarize it. A summary is the log’s reading of a source; a quote is the source. Only the second survives the summarizer being wrong, and being wrong is the normal case — a paraphrase that drops a hedge, promotes a subsample to a headline, or converts “we cannot reject” into “no effect” is indistinguishable from an accurate one once the tab is closed. So the standing requirement is that a factual claim carries the words it came from.

What must be quoted. Anything a reader might later want to check or dispute: headline results and their qualifiers, sample sizes and denominators, the authors’ own statements of scope and limitation, and any wording this log leans on. Verbatim disclaimers matter as much as verbatim findings, and often more, because the disclaimer is what gets lost first when a figure is passed along. Where a source hedges, quote the hedge rather than the hedged claim.

What cannot be quoted, and must say so. Some claims have no source sentence behind them: statistics computed here from a data file, values read off a chart, counts assembled across entries, comparisons this log is making rather than reporting. These are the log’s own assertions, not the source’s, and they should be written as such — “computed from the CSV”, “read from the figure”, “this log’s comparison” — so that the absence of a quote reads as a category difference rather than as laziness.

Where the quote goes. A quote that is self-contained and short enough to read in passing belongs in the sentence making the claim. A quote that needs surrounding context to be fair, or that runs past roughly one line, goes in a footnote, so the claim stays readable and the evidence stays adjacent.1 Quoting more than the claim strictly needs is a venial sin here; quoting less than makes the claim checkable is not.

1 This is a footnote — the container for a quote that is worth preserving in full but would otherwise swamp the bullet that depends on it.

Verbatim discipline. Reproduce the source’s wording, spelling, and emphasis. Mark every omission with an ellipsis and every insertion with square brackets. Never silently repair grammar, expand an abbreviation, or convert a percentage into a ratio inside quotation marks. Attribute the quote to a specific document, not to an organization, when the same body has said different things at different dates. If a quote comes from an abstract, a summary page, or press coverage rather than the body of the work, the status line says so — an abstract is a source’s advertisement for itself.

A figure is a claim, so say where it came from. Most figures here are reproduced from the source and are covered by the entry’s status line — but reproducing one means the figure itself was read, which is more than verified-abstract normally implies, so those status lines say which figure. Five figures are instead generated by tools/sources_figures.py and are the log’s own constructions rather than anyone’s: the AIxCC comparison and the modded-nanogpt curve, plotted from numbers recorded in their own entries; the chronology timeline, parsed from the chronology table so it cannot drift from it; the efficiency-rate comparison, fitted to a vendored cost-curve CSV; and the exponent-bound series, extracted from the exponent database by tools/antedb_extract.py and vendored as a CSV [→ ANTEDB rates]. If a generated figure needs to change, change the data or the script and regenerate — never edit the image.

The log may hold more than the argument uses, but it must say so. A quarry contains more stone than the building. An entry the argument has not yet drawn on carries a **Unused:** line saying so and naming where it would go, and the validator counts those separately rather than failing. The line is not a parking space for weak material: an entry still has to meet every other requirement to be here at all. When the argument starts citing an entry, delete the line — the validator fails if a cited entry is still marked unused, which keeps the two files honest about what has actually been used.

Date everything, and keep three kinds of date apart. Each entry carries a **Dates:** line distinguishing when the evidence was produced (the run, the competition, the window a study covers), when it was published (and every subsequent version, because arXiv figures move), and when this log last checked it. These come apart constantly and the difference changes what a source can support: a benchmark posted in 2025 and revised in 2026 cannot be quoted from memory, a vendor claim and a reproduction two days apart bound how much reproduction was possible, and a 2024 measurement describes tooling nobody uses now. Where a date cannot be established, the line says so rather than omitting it — an unknown date is itself a finding about the source. When a date here changes, update the chronology at the foot of the document in the same edit.

Known gaps are part of the log. Questions with thin evidence are listed at the end. Silence about a gap reads as absence of an effect, which is usually wrong.

An index of every entry

Generated from the entries by tools/sources_index.py, so it cannot drift from them. “Source” is who produced the work, read from each entry’s italic type line; “Checked” is how far this log has verified it, read from the status line; “Used” says whether the companion argument cites the entry or the log is holding it. The two columns worth reading together are Source and Checked, because they answer different questions and the second is often mistaken for the first.

120 entries. By how far they have been checked: 54 verified, 41 verified-abstract, 17 unverified, 8 derived. By who produced them: 75 independent, 29 vendor, 8 derived, 4 government, 2 mixed, 1 disavowed, 1 industry. 19 carry a vendor self-report somewhere in their figures, and 79 are held unused by the argument.

Entry Year Domain Source Checked Bears on Used
Our World in Data: cost curves for 66 technologies 2016 Aggregate measures independent verified Q1 Q8 cited
Bloom, Jones, Van Reenen, and Webb: are ideas getting harder to find? 2020 Aggregate measures independent verified Q1 Q5 held
METR: task-completion time horizons 2025 Aggregate measures independent verified Q1 Q6 Q8 held
Analytic Number Theory Exponent Database, ANTEDB 2025 Aggregate measures independent verified Q1 Q7 Q8 cited
Comparing efficiency rates across domains 2026 Aggregate measures derived derived Q1 Q5 Q7 held
Epoch AI: ML hardware price-performance 2023 Aggregate measures independent verified Q1 Q5 Q6 held
Hao and co-authors: AI expands individual impact and contracts collective focus 2024 Aggregate measures independent verified-abstract Q1 Q4 Q7 held
Acemoglu: the simple macroeconomics of AI 2024 Aggregate measures independent verified Q1 Q5 Q7 Q8 held
Aghion and Bunel: a higher estimate, with the ideas channel explicitly omitted 2024 Aggregate measures independent verified Q1 Q5 Q6 Q8 held
Toner-Rodgers, “AI, Scientific Discovery, and Product Innovation” — WITHDRAWN, do not cite 2024 Aggregate measures disavowed verified Q4 cited
Benjamin Jones: AI in R&D 2025 Conceptual models independent verified Q2 Q3 Q5 Q7 cited
Bazzichi, Riccaboni, and Castellacci: recombinant innovation 2026 Conceptual models independent verified Q3 Q5 Q7 cited
DeepMind: conjecture machines and the validation bottleneck 2026 Conceptual models vendor verified, vendor Q3 Q7 cited
Davidson, Halperin, Houlden, and Korinek: recursive R&D feedback 2026 Conceptual models independent verified Q1 Q6 cited
Aghion, Jones, and Jones: AI in the idea production function 2017 Conceptual models independent verified Q1 Q2 Q5 Q6 held
Korinek and Suh: whether wages collapse depends on the tail of task complexity 2024 Conceptual models independent verified-abstract Q2 Q3 Q5 Q6 held
Ide and Talamas: autonomous AI helps the most knowledgeable, assistive AI helps the least 2023 Conceptual models independent verified-abstract Q2 Q3 Q4 held
Besiroglu, Emery-Xu, and Thompson: AI-augmented R&D is more capital-intensive 2022 Conceptual models independent verified-abstract Q1 Q5 Q6 held
UK AISI: multi-step cyber attack scenarios 2026 Cyber government verified Q1 Q2 Q5 Q6 Q7 cited
Google Big Sleep 2025 Cyber vendor verified, vendor Q2 Q4 cited
DARPA AIxCC finals 2025 Cyber government verified Q1 Q2 Q5 Q6 Q8 cited
CyberGym 2025 Cyber independent verified-abstract Q2 Q5 Q7 cited
Anthropic Mythos preview 2026 Cyber vendor unverified, vendor Q1 Q7 Q8 cited
XBOW on HackerOne 2025 Cyber vendor unverified, vendor Q3 Q5 Q7 cited
curl ends its bug bounty 2026 Cyber independent verified Q3 Q5 cited
Gemini and developer experience: security of the resulting code 2026 Cyber independent verified-abstract Q3 Q4 held
Cyber labour-market indicators 2026 Cyber industry unverified Q3 Q4 held
Vidoc Security reproduction 2026 Cyber independent unverified Q1 Q4 cited
AISLE: the jagged frontier 2026 Cyber independent verified Q7 Q8 cited
Palo Alto Networks Mythos deployment 2026 Cyber vendor unverified Q1 Q8 cited
UK AISI Mythos runs, via dbreunig 2026 Cyber independent verified Q4 Q5 Q8 cited
Lyptus: offensive-cyber time horizons 2026 Cyber independent verified Q6 Q8 cited
ARTEMIS pentest study 2025 Cyber independent verified-abstract Q3 Q8 cited
DARPA Cyber Grand Challenge: the pre-LLM autonomy baseline 2016 Cyber government verified Q1 Q2 Q4 held
Google Project Zero, “Project Naptime”: tooling versus the model 2024 Cyber independent verified Q1 Q2 Q6 Q8 held
Fang and co-authors: autonomous exploitation of one-day vulnerabilities 2024 Cyber independent verified-abstract Q2 Q4 Q7 held
Fang and co-authors: teams of agents on zero-day vulnerabilities 2024 Cyber independent verified-abstract Q2 Q5 Q8 held
Cybench: CTF tasks with human solve times attached 2024 Cyber independent verified-abstract Q1 Q2 Q4 Q8 held
Meta CYBERSECEVAL 3: a lab reporting a null on its own model 2024 Cyber vendor verified, vendor Q2 Q7 held
HackerOne: platform statistics on agent-submitted reports 2025 Cyber vendor unverified, vendor Q2 Q3 Q7 held
NIST on record CVE growth and the enrichment backlog 2026 Cyber government verified Q1 Q3 Q7 held
OpenAI: Erdős unit-distance disproof 2026 Math vendor verified Q1 Q2 Q3 Q7 cited
Erdős #728 2026 Math independent verified-abstract Q2 Q4 cited
Terence Tao commentary 2025–2026 Math independent verified Q1 Q5 Q7 cited
AlphaProof Nexus: formal proof search on open problems 2026 Math vendor verified-abstract Q1 Q2 Q4 Q5 Q7 Q8 held
FunSearch: cap sets and bin packing 2023 Math vendor verified-abstract, vendor Q1 Q2 Q7 held
AlphaProof and IMO 2024 silver 2024 Math vendor verified-abstract, vendor Q6 Q8 held
Erdős-problems wiki: AI contributions 2026 Math independent verified Q4 Q7 Q8 cited
GPT-5 literature retrieval 2025 Math mixed verified Q4 Q8 cited
IMO 2025 gold 2025 Math vendor verified Q8 cited
Epoch AI: FrontierMath 2026 Math independent verified Q6 Q8 cited
Kissing number: AlphaEvolve then a human 2025 Math independent verified Q1 Q3 cited
AlphaGeometry: olympiad geometry from synthetic data 2024 Math vendor verified, vendor Q1 Q2 Q7 Q8 held
AlphaGeometry 2: past the average gold medalist 2025 Math vendor verified-abstract, vendor Q1 Q4 Q7 Q8 held
PutnamBench: a formal benchmark that started near the floor 2024 Math independent verified-abstract Q1 Q2 Q7 Q8 held
The Equational Theories Project: 22 million implications, formally settled 2024–2025 Math independent verified-abstract Q2 Q3 Q4 Q7 held
Aristotle: formally verified IMO 2025 proofs 2025 Math vendor unverified, vendor Q2 Q4 Q8 held
Ringer in Nature: mathematicians’ hands-on assessment of AlphaProof 2025 Math independent verified Q2 Q3 Q4 Q7 Q8 held
AlphaEvolve across 67 mathematical problems, including where it failed 2025 Math mixed unverified, vendor Q1 Q2 Q4 Q7 Q8 held
ThetaEvolve: an 8B open model passes AlphaEvolve’s bounds 2025 Math independent unverified Q2 Q4 Q5 Q6 held
HorizonMath: unsolved problems with cheap verification 2026 Math independent unverified Q2 Q7 Q8 held
Williams: what played to AI’s strengths in the unit-distance disproof 2026 Math independent verified Q3 Q4 Q7 held
Sherry and Thompson: how fast do algorithms improve? 2021 Algorithms independent verified Q1 Q5 Q6 Q7 held
Bixby: LP and mixed-integer programming solver speedups 2012 Algorithms vendor verified Q1 Q5 Q6 Q7 held
Grace: algorithmic progress in six domains 2013 Algorithms independent verified Q1 Q7 Q8 held
Epoch AI: algorithmic progress in language models 2024 Algorithms independent verified-abstract Q1 Q8 cited
Hernandez and Brown: measuring the algorithmic efficiency of neural networks 2020 Algorithms independent verified Q1 Q5 held
Erdil and Besiroglu: algorithmic progress in computer vision 2022 Algorithms independent verified Q1 Q5 Q7 held
The SAT Museum: thirty years of solvers on one machine 2023 Algorithms independent verified Q1 Q4 Q6 Q7 held
Epoch AI: LLM inference price declines 2025 Algorithms independent verified Q1 Q5 Q6 held
AlphaEvolve 2025 Algorithms vendor verified-abstract, vendor Q1 Q2 Q7 cited
TTT-Discover 2026 Algorithms independent unverified Q2 Q4 Q6 cited
modded-nanogpt speedrun 2024–2026 Algorithms independent verified Q1 Q2 Q3 cited
Karpathy autoresearch 2026 Algorithms independent unverified Q1 Q4 cited
AlphaTensor and AlphaDev: pre-LLM algorithm discovery 2022–2023 Algorithms vendor verified-abstract, vendor Q1 Q2 Q7 held
Cui and co-authors: pooled coding-assistant RCTs 2025 Algorithms independent unverified Q1 Q3 Q4 Q7 held
RE-Bench 2024 Algorithms independent verified Q3 Q5 Q8 cited
METR: RCT on experienced open-source developers 2025 Algorithms independent verified-abstract Q1 Q3 Q4 Q7 cited
AlgoTune 2025 Algorithms independent verified-abstract Q1 Q7 Q8 cited
MLGym 2025 Algorithms independent verified-abstract Q1 cited
SWE-fficiency 2025 Algorithms independent verified-abstract Q7 Q8 cited
GSO 2025 Algorithms independent verified-abstract Q7 Q8 cited
PERFOPT-Bench relay pilot 2026 Algorithms independent verified-abstract Q5 Q6 Q8 cited
Performance-benchmark reliability audit 2026 Algorithms independent verified-abstract Q8 cited
SWE-bench: the benchmark that dated the floor at 1.96% 2023 Algorithms independent verified Q1 Q2 Q8 held
MLE-bench: agents on Kaggle competitions against public leaderboards 2024 Algorithms vendor verified-abstract Q2 Q4 Q5 Q8 held
PaperBench: replicating ICML papers from scratch 2025 Algorithms vendor verified-abstract Q2 Q3 Q4 Q8 held
CORE-Bench: computational reproducibility as the floor task 2024 Algorithms independent verified-abstract Q2 Q4 Q8 held
The AI Scientist: automated papers at under fifteen dollars each 2024 Algorithms vendor verified-abstract, vendor Q2 Q4 Q5 Q8 held
The AI Scientist-v2: one autonomous manuscript through real peer review 2025 Algorithms vendor verified-abstract, vendor Q1 Q2 Q4 Q8 held
Peng and co-authors: the Copilot randomized trial 2023 Algorithms vendor verified Q1 Q3 Q4 Q7 Q8 held
Hoffmann and co-authors: Copilot shifts what developers do 2024 Algorithms vendor verified Q2 Q3 Q4 Q6 held
Song and co-authors: Copilot raises contributions and coordination cost together 2024 Algorithms independent verified-abstract Q1 Q3 Q4 Q7 held
DORA: throughput up, delivery stability still down 2025 Algorithms vendor verified, vendor Q1 Q3 Q8 held
GitClear: code duplication up, refactoring down 2025 Algorithms vendor verified, vendor Q1 Q3 Q8 held
Kapoor and co-authors: AI agents that matter 2024 Measurement and benchmark validity independent verified-abstract Q1 Q5 Q8 held
The SWE-bench illusion: memorization rather than reasoning 2025 Measurement and benchmark validity independent verified Q1 Q7 Q8 held
Sakana on kernel benchmarks: a vendor conceding exploitable loopholes 2025 Measurement and benchmark validity vendor verified-abstract Q1 Q2 Q8 held
GDPval: a benchmark built to predict economic value, with its authors’ limits 2025 Measurement and benchmark validity vendor verified Q1 Q2 Q8 held
AI-discovered drugs: the proxy clears, the objective does not 2024 Measurement and benchmark validity vendor verified-abstract Q1 Q2 Q5 Q8 held
Stack Overflow developer survey: adoption high, trust low 2025 Measurement and benchmark validity independent unverified Q1 Q3 Q8 held
Simple baselines match code evolution, so the machinery may not be what works 2026 Measurement and benchmark validity independent unverified Q1 Q4 Q8 held
Dell’Acqua and co-authors: the jagged technological frontier 2023 Expertise and demand independent verified-abstract Q4 Q5 Q7 held
Brynjolfsson, Li, and Raymond: generative AI at work 2023 Expertise and demand independent unverified Q3 Q4 held
Otis and co-authors: the uneven impact on Kenyan entrepreneurs 2024 Expertise and demand independent verified Q4 Q7 held
Does generative AI narrow education-based productivity gaps? 2026 Expertise and demand independent verified-abstract Q4 held
Doshi and Hauser: AI raises individual creativity and lowers collective diversity 2024 Expertise and demand independent unverified Q5 Q7 held
Noy and Zhang: ChatGPT compresses the writing productivity distribution 2023 Expertise and demand independent verified-abstract Q1 Q4 Q8 held
Hui, Reshef, and Zhou: freelance demand fell, and top freelancers fell hardest 2023 Expertise and demand independent verified Q3 Q4 Q7 held
Demirci, Hannane, and Zhu: postings for automation-prone freelance work fell 21% 2024 Expertise and demand independent verified-abstract Q3 Q5 Q7 held
Brynjolfsson, Chandar, and Chen: entry-level employment in AI-exposed occupations 2025 Expertise and demand independent verified, vendor Q3 Q4 Q7 held
Aghion and co-authors: French firms that adopted AI grew 2025 Expertise and demand independent verified-abstract Q3 Q5 Q7 held
Zhao and co-authors: AlphaFold barely changed who collaborates 2025 Expertise and demand independent verified Q3 Q4 Q7 held
How fast analytic number theory’s exponents actually improve 2026 Syntheses assembled here derived derived Q1 Q6 Q7 Q8 held
Inventory of the 67 AlphaEvolve problems, as the frame for a historical baseline 2026 Syntheses assembled here derived derived Q1 Q7 Q8 held
AI record steps against human steps on the same quantities 2026 Syntheses assembled here derived derived Q1 Q4 Q5 Q7 held
Every rate in the log that has a denominator 2026 Syntheses assembled here derived derived Q1 Q2 Q5 Q7 Q8 held
Every cost-per-result figure in the log 2026 Syntheses assembled here derived derived Q1 Q5 Q6 Q8 held
What “autonomous” turned out to mean, case by case 2026 Syntheses assembled here derived derived Q2 Q4 Q8 held
The randomized and quasi-experimental evidence, in one place 2026 Syntheses assembled here derived derived Q1 Q3 Q4 Q7 held

Which entries bear on which question

Also generated. The row lengths are the point: Q1 and Q7 have more evidence than anyone can synthesize, and the rows that matter most for distinguishing the theories are among the shortest. Where a row is short, read the corresponding paragraph in “Known gaps” before concluding anything from it.

Question Entries bearing on it
Q1 growth rate (66)
Does AI raise the rate of efficiency growth, and by how much?
owid-tech-costs · ideas-harder · metr-horizons · antedb · efficiency-rates · hardware-price-performance · ai-science-bibliometrics · acemoglu-simple · aghion-bunel · recursive-rd · aghion-jones-jones · besiroglu-airnd · aisi-cyber · aixcc · mythos · vidoc · palo-alto · cyber-grand-challenge · naptime · cybench · nvd-growth · unit-distance · tao · alphaproof-nexus · funsearch · kissing · alphageometry · alphageometry2 · putnambench · alphaevolve-math · sherry-thompson · bixby · grace-six-domains · algo-progress · alexnet-efficiency · vision-progress · sat-museum · inference-prices · alphaevolve · nanogpt · autoresearch · alphatensor-alphadev · copilot-rcts · metr-dev-rct · algotune · mlgym · swebench · ai-scientist-v2 · copilot-rct · copilot-opensource · dora · gitclear · agents-that-matter · swebench-illusion · robust-kbench · gdpval · ai-drugs · stackoverflow · simple-baselines · noy-zhang · antedb-rates · alphaevolve-inventory · alphaevolve-records · denominators · cost-per-result · experimental-evidence
Q2 autonomy (43)
Can AI produce a complete result without a human driving it?
jones-aird · aghion-jones-jones · korinek-suh · ide-talamas · aisi-cyber · big-sleep · aixcc · cybergym · cyber-grand-challenge · naptime · fang-oneday · fang-teams · cybench · cyberseceval3 · hackerone · unit-distance · erdos-728 · alphaproof-nexus · funsearch · alphageometry · putnambench · equational-theories · aristotle · ringer-alphaproof · alphaevolve-math · thetaevolve · horizonmath · alphaevolve · ttt-discover · nanogpt · alphatensor-alphadev · swebench · mlebench · paperbench · corebench · ai-scientist · ai-scientist-v2 · copilot-composition · robust-kbench · gdpval · ai-drugs · denominators · autonomy-ladder
Q3 demand (35)
What happens to demand for human researchers?
jones-aird · recombinant-innovation · validation-bottleneck · korinek-suh · ide-talamas · xbow · curl · gemini-dev-security · cyber-labor · artemis · hackerone · nvd-growth · unit-distance · kissing · equational-theories · ringer-alphaproof · williams-context · nanogpt · copilot-rcts · rebench · metr-dev-rct · paperbench · copilot-rct · copilot-composition · copilot-opensource · dora · gitclear · stackoverflow · support-agents · freelancer-demand · posting-demand · canaries · french-firms · alphafold-collaboration · experimental-evidence
Q4 expertise (47)
How expert are the people who make discoveries with AI?
ai-science-bibliometrics · toner-rodgers · ide-talamas · big-sleep · gemini-dev-security · cyber-labor · vidoc · aisi-runs · cyber-grand-challenge · fang-oneday · cybench · erdos-728 · alphaproof-nexus · erdos-wiki · gpt5-retrieval · alphageometry2 · equational-theories · aristotle · ringer-alphaproof · alphaevolve-math · thetaevolve · williams-context · sat-museum · ttt-discover · autoresearch · copilot-rcts · metr-dev-rct · mlebench · paperbench · corebench · ai-scientist · ai-scientist-v2 · copilot-rct · copilot-composition · copilot-opensource · simple-baselines · jagged-frontier · support-agents · kenya-entrepreneurs · education-gap-rct · noy-zhang · freelancer-demand · canaries · alphafold-collaboration · alphaevolve-records · autonomy-ladder · experimental-evidence
Q5 returns (38)
How fast do returns to AI spend diminish against human labor?
ideas-harder · efficiency-rates · hardware-price-performance · acemoglu-simple · aghion-bunel · jones-aird · recombinant-innovation · aghion-jones-jones · korinek-suh · besiroglu-airnd · aisi-cyber · aixcc · cybergym · xbow · curl · aisi-runs · fang-teams · tao · alphaproof-nexus · thetaevolve · sherry-thompson · bixby · alexnet-efficiency · vision-progress · inference-prices · rebench · perfopt · mlebench · ai-scientist · agents-that-matter · ai-drugs · jagged-frontier · idea-diversity · posting-demand · french-firms · alphaevolve-records · denominators · cost-per-result
Q6 intertemporal (23)
Is it better to spend on AI now or later?
metr-horizons · hardware-price-performance · aghion-bunel · recursive-rd · aghion-jones-jones · korinek-suh · besiroglu-airnd · aisi-cyber · aixcc · lyptus · naptime · alphaproof-imo · frontiermath · thetaevolve · sherry-thompson · bixby · sat-museum · inference-prices · ttt-discover · perfopt · copilot-composition · antedb-rates · cost-per-result
Q7 incidence (57)
Which fields and starting points does AI affect most?
antedb · efficiency-rates · ai-science-bibliometrics · acemoglu-simple · jones-aird · recombinant-innovation · validation-bottleneck · aisi-cyber · cybergym · mythos · xbow · aisle · fang-oneday · cyberseceval3 · hackerone · nvd-growth · unit-distance · tao · alphaproof-nexus · funsearch · erdos-wiki · alphageometry · alphageometry2 · putnambench · equational-theories · ringer-alphaproof · alphaevolve-math · horizonmath · williams-context · sherry-thompson · bixby · grace-six-domains · vision-progress · sat-museum · alphaevolve · alphatensor-alphadev · copilot-rcts · metr-dev-rct · algotune · swefficiency · gso · copilot-rct · copilot-opensource · swebench-illusion · jagged-frontier · kenya-entrepreneurs · idea-diversity · freelancer-demand · posting-demand · canaries · french-firms · alphafold-collaboration · antedb-rates · alphaevolve-inventory · alphaevolve-records · denominators · experimental-evidence
Q8 benchmarks (58)
Which benchmarks predict real value?
owid-tech-costs · metr-horizons · antedb · acemoglu-simple · aghion-bunel · aixcc · mythos · aisle · palo-alto · aisi-runs · lyptus · artemis · naptime · fang-teams · cybench · alphaproof-nexus · alphaproof-imo · erdos-wiki · gpt5-retrieval · imo · frontiermath · alphageometry · alphageometry2 · putnambench · aristotle · ringer-alphaproof · alphaevolve-math · horizonmath · grace-six-domains · algo-progress · rebench · algotune · swefficiency · gso · perfopt · perf-reliability · swebench · mlebench · paperbench · corebench · ai-scientist · ai-scientist-v2 · copilot-rct · dora · gitclear · agents-that-matter · swebench-illusion · robust-kbench · gdpval · ai-drugs · stackoverflow · simple-baselines · noy-zhang · antedb-rates · alphaevolve-inventory · denominators · cost-per-result · autonomy-ladder

Aggregate measures

Our World in Data: cost curves for 66 technologies (2016)

Independent (data aggregator, from an academic dataset). Our World in Data’s “The cost of 66 different technologies over time” plots unit cost against year on a log axis for 66 technologies, drawn from the Santa Fe Performance Curve Database as compiled by Farmer and Lafond (2016), “How predictable is technological progress?” This is the closest available picture of efficiency progress — cost per unit of output — measured consistently across many domains at once. Figures below are computed from the chart’s CSV download, not read off the image.

Log-scale cost curves for dozens of technologies, most falling as near-straight lines over decades.

Unit cost against year for a selection of the 66 series, on a log axis.

Near-straight lines on a log axis, at slopes differing by about an order of magnitude between domains. The series are unit cost against year, and all of them end before AI played any part in them.

Note on quotation. Almost every figure in this entry is computed here from the chart’s CSV download rather than read from a sentence, so there is little to quote and the numbers are this log’s arithmetic. They are reproducible from the file; the two quotations below are the chart’s own framing.

  • Steady exponential improvement is the normal case, and 64 of the 66 series end cheaper than they began. Computed from the CSV. Only two end more expensive: nuclear electricity (1970–1989, $0.26 to $3.37, about +14.4% a year) and crude oil (1946–1968).
  • The rates differ by an order of magnitude across domains. Annualized cost change over each series’ own span, computed here: DNA sequencing −56.8% (2001–2013), hard disk drive −43.8% (1988–2007), transistor −39.2% (1968–2005), DRAM −36.0% (1971–2007), photovoltaics −9.6% (1980–2013), Danish wind turbines −3.5% (1981–2000).
  • On a log axis most series are close to straight lines. That is the empirical content of Wright’s and Moore’s laws: within a domain the improvement rate is roughly constant for decades, so a domain’s rate — its slope — is the natural thing for a new technology to change. Read from the chart, and the framing is this log’s.
  • Units are not comparable across series, and the chart says so. Its subtitle: “The cost of each technology is expressed in different units, chosen for visualization purposes.” Its note: “DNA sequencing is divided by 1000 to fit on the chart.” Only within-series rates of change are meaningful, never levels between series.
  • The data stops in 2013. The series run 1929–2013, so this dataset cannot speak to AI’s effect at all. It establishes the shape of the outcome variable and the pre-AI baseline, nothing more.
  • Provenance. Our World in Data describes the underlying data as “adapted from Farmer and Lafond,” from the Santa Fe Performance Curve Database, published as “How predictable is technological progress?”
  • Dates: underlying series 1929–2013; Farmer and Lafond published 2016; CSV retrieved 2026-07-26.
  • Bears on: Q1 growth rate, Q8 benchmarks — it defines the shape of the outcome variable.
  • Links: chart · Technological Change topic page · Farmer and Lafond (2016)
  • Status: verified against the chart’s CSV download (1,256 rows, 66 entities, 1929–2013), retrieved 2026-07-26. Chart snapshot stored locally at posts/images/owid-tech-cost-curves.png.

Bloom, Jones, Van Reenen, and Webb: are ideas getting harder to find? (2020)

Independent (academic, peer-reviewed). The canonical measurement of the fishing-out term. Research productivity — ideas produced per researcher — is estimated across US aggregate data, semiconductors, agricultural crop yields, medical innovation, and firm-level panels. This is the entry that calibrates \(\beta\) in the research production function the argument is built on, and it is the pre-AI baseline against which any claimed AI effect has to be judged.

Two lines on a log scale from the 1930s to the 2000s: research productivity falling from 1 to about 1/40, and the effective number of researchers rising from 1 to about 25.

US aggregate research productivity against the effective number of researchers, 1930s to 2000s, both indexed to 1 in the 1930s.

The paper’s Figure 2. Both series are indexed to 1 in the 1930s and plotted on log scales: the effective number of researchers rises by a factor of 23, and research productivity — ideas per researcher — falls by a factor of 41. Research productivity here is the ratio of TFP growth to research effort, so the two plotted lines are the ratio and its denominator.

  • Aggregate research productivity halves about every 13 years. In the authors’ words, “Taking the US aggregate number as representative, research productivity falls in half every 13 years: ideas are getting harder and harder to find.” The implication they draw is the one that matters for this project: “just to sustain constant growth in GDP per person, the United States must double the amount of research effort every 13 years to offset the increased difficulty of finding new ideas.”
  • Over the long run, effort rose by 23× and productivity fell by 41×. “Since the 1930s, research effort has risen by a factor of 23, an average growth rate of 4.3 percent per year. Research productivity has fallen by an even larger amount, by a factor of 41.” An earlier version of this entry said productivity fell by “a comparable factor” to the rise in effort, which understates it; the two figures are 23 and 41.
  • Moore’s Law is the cleanest case, and it is stark. “The number of researchers required today to achieve the famous doubling of computer chip density is more than 18 times larger than the number required in the early 1970s.” The productivity decline follows arithmetically: “Assuming a constant growth rate for Moore’s Law, the implication is that research productivity has fallen by this same factor of 18, an average rate of 6.8 percent per year.” Semiconductors are nonetheless the least diminishing sector they study — “A is growing at 35 percent per year, while research productivity is falling at 7 percent per year,” which they read as “the sector with the least degree of diminishing returns in idea production.”
  • The decline is not an artifact of one sector. “Our robust finding is that research productivity is falling sharply everywhere we look,” across crop yields, medical innovation, and firm-level panels as well as semiconductors. The paper opens by calling it settled: “A first-order fact of growth empirics is that research productivity is falling sharply.”
  • The paper is explicit that ideas-per-dollar is the wrong measure, and this log contains a lot of ideas-per-dollar. The authors note that “A large literature documents that the flow of new ideas per research dollar is declining,” and then disqualify it: “essentially all the idea-driven growth models in the literature predict that ideas per (real) research dollar will be declining… In other words, these natural measures are not really informative about whether research faces constant or diminishing returns.” The theory-relevant object is ideas per researcher. Several entries here report cost per result — AIxCC’s $152 per task, AlphaProof Nexus’s few hundred dollars per problem, AISI’s $12,500 per attempt — and this is the warning that none of them, on its own, measures returns.
  • Why it is load-bearing. Every domain in this log has a large and rising research-effort denominator that mostly goes unmeasured. A falling yield per attempt is the normal state of research, so an apple-picking prediction of falling yield is only distinctive if it falls faster than this baseline. Nothing in this log yet makes that comparison — this sentence is the log’s own assessment, not the paper’s.
  • Dates: NBER working paper 23782 issued 2017-09-08; published in the American Economic Review 110(4), April 2020; the underlying series run from the 1930s to about 2015.
  • Bears on: Q1 growth rate, Q5 returns.
  • Unused: not yet cited in the argument. It belongs wherever the formalization introduces \(\beta\), as the empirical anchor for the fishing-out exponent and as the reason a declining yield curve is not by itself evidence for apple-picking.
  • Links: AER · NBER w23782 · working paper PDF
  • Status: verified against the AER abstract and the working-paper PDF, retrieved 2026-07-26. The 18× and 6.8% figures are quoted from the paper; the aggregate factor is read from its Figure 2 discussion rather than a table. Figure reproduced from the source.

METR: task-completion time horizons (2025)

Independent (METR). The parent methodology for every “time horizon” number in this log, including the offensive-cyber series. Frontier agents are scored on tasks whose difficulty is measured by how long human experts take, and the reported metric is the human task length at which an agent succeeds half the time.

Log-scale plot of task length at 50 percent success against model release date, rising in a straight line from GPT-2 in 2019 to o3 in 2025.

The 50% task-completion time horizon against model release date.

A straight line on a log axis over six years, which is the same functional form as the technology cost curves above. What it does not show is what a task is worth: a horizon measures length at fixed reliability, not depth.

  • The metric, in the authors’ definition. They “propose a new metric: 50%-task-completion time horizon, the time humans typically take to complete tasks that AI models can complete with 50% success rate.” Note the direction of the measurement: it is a statement about human task length, not about how long the agent runs.
  • The 50% horizon doubled about every seven months from 2019 to 2025. “Frontier AI time horizon has doubled approximately every seven months since 2019, though the trend may have accelerated since 2024.” The original release put frontier models such as o3 at a 50% horizon of roughly 110 minutes, with a doubling time of 195.8 days [162, 223].
  • The updated task suite left the long trend intact but shortened the recent one. Time Horizon 1.1 expanded the suite from 170 to 228 tasks, and tasks of 8 hours or more from 14 to 31. On the full period, “This hybrid trend shows exactly the same doubling time as the TH1 trend, of 196 days (7 months).” On the recent period it moves: “We also report below the doubling time since 2024: this was at 109 days under TH1, and falls to 89 days under TH1.1.” From 2023 onward the estimate falls from 165.3 days to 130.8 days [107, 161].
  • The acceleration is real but measured on thin data. Only 5 of the 8-hour-plus tasks are human-baselined in the updated suite, which is the binding constraint on estimating a horizon approaching a day. METR also cautions against reading the confidence intervals as a stability guarantee, since “those confidence intervals represent the likelihood of getting the same estimate with an entirely new set of tasks, while in fact there is substantial overlap in the tasks contained in TH1.”
  • The authors’ own limitation, quoted because it is the one usually dropped. “Note that this plot does not account for future changes in the trend or external validity concerns, which are responsible for the majority of our uncertainty.” Their account of the mechanism is similarly hedged: the rise “seems to be primarily driven by greater reliability, ability to adapt to mistakes, logical reasoning, and capacity for tool use.”
  • The trend appears in other domains too. METR reports analyzing “9 benchmarks for scientific reasoning, math, robotics, computer use, and self-driving in terms of time-horizon trends; we observe generally similar rates of improvement to the 7-month doubling time in our original time-horizon work.”
  • Why it belongs here, and the caveat that limits it. It is the closest thing to a general dated capability series at fixed scaffold, which is the measurement Q6 needs and mostly lacks. It is also the methodological parent of the Lyptus cyber horizons [→ Lyptus], so the two are not independent confirmations of an acceleration. And a horizon says nothing about what a result is worth: it cannot distinguish a long shallow task from a long deep one, which is the distinction this project turns on. Both points are this log’s, not METR’s.
  • Dates: arXiv 2025-03-18 (v4 2026-07-10); METR write-up 2025-03-19; Time Horizon 1.1 published 2026-01-29; the model series runs 2019 to 2026.
  • Bears on: Q1 growth rate, Q6 intertemporal, Q8 benchmarks.
  • Unused: not yet cited in the argument. The natural place is the Q6 discussion, where the argument currently relies on Lyptus alone and would be stronger noting that Lyptus inherits this methodology.
  • Links: METR time horizons · arXiv 2503.14499 · original write-up · Time Horizon 1.1
  • Status: verified against the arXiv abstract and the Time Horizon 1.1 comparison table, retrieved 2026-07-26.

Analytic Number Theory Exponent Database, ANTEDB (2025)

Independent (academic database). Tao, Trudgian, and Yang’s ANTEDB systematically records known theorems, conjectures, and relationships for exponents appearing in analytic number theory: exponent pairs, exponential sum bounds, zero-density and moment bounds for the Riemann zeta function, large value estimates, additive energy bounds, and exponents governing prime distributions. It is the closest thing research mathematics has to a per-problem efficiency curve.

  • Bounds are a continuous outcome variable; solved-or-not is a binary one. Each exponent is a number whose best known value is recorded with a date and a proof, so progress on it is a monotone dated series rather than a flag that flips once. This is what makes bounds, rather than problem counts, the natural analogue of a cost curve — this log’s argument for using the database, not a claim the database makes.
  • That series has now been extracted and plotted. The database’s Python codebase can restrict the literature to results published up to a given year and then solve for the best bound derivable from that restriction, which turns the database into a dated efficiency curve. Six such series are extracted and plotted in a derived entry below [→ ANTEDB rates]; their halving times run from 82 to 1,204 years. The database’s own figures plot exponent-pair regions in parameter space rather than history, so the time series appears to be new here.
  • The database is both human-readable and executable. Tao: “Information on these exponents is collected both in a LaTeX ‘blueprint’ that is available as a human-readable set of web pages, and as part of our Python codebase.” Formalization is aspirational rather than done: “In the future one could also imagine the data being collected in a Lean formalization, but at present the database only contains a placeholder Lean folder.”
  • Automated search over the recorded relations already improved the state of the art. The launch paper reports “four new exponent pairs; several new zero density estimates; and new estimates on the additive energy of zeroes of the Riemann zeta function.” Tao describes how, and the parenthesis is the important part.2
  • That automation is optimization over a relation database, not an LLM agent. The improvement came from systematically combining bounds already in the literature. It is evidence about how much unpicked fruit sits in an uncollated literature, not about model capability. This distinction is this log’s, and it is the reason the entry sits in aggregate measures rather than in the math section.
  • It is a living database, not a benchmark. It has no fixed problem set, no scoring rule, and invites contributions — “We are hoping that the ANTEDB will receive more contributions in the future, for instance expanding to other types of exponents, or to update the database as new results are obtained (or old ones added)” — so it cannot be used as a capability time series without imposing structure it does not have.
  • The stated purpose is collation, and the hoped-for payoff is exactly the apple-picking mechanism. “By providing a centralized database, it aids in the verification and progression of results in the field, potentially leading to breakthroughs or simplifications in proofs.” A literature whose existing results have not been systematically combined is a tree with reachable fruit on it.
  • It builds on an earlier static table. The precursor is Trudgian and Yang’s “Toward optimal exponent pairs.”
  • Dates: precursor exponent-pair tables arXiv 2023-06-09 (v3 2024-07-15); Tao’s launch post 2025-01-28; blueprint PDF dated 2026-07-03; retrieved 2026-07-26.
  • Bears on: Q1 growth rate, Q7 incidence, Q8 benchmarks.
  • Links: repo · blueprint site · launch paper, arXiv 2501.16779 · Tao’s launch post · Trudgian and Yang precursor
  • Status: verified against the repository README, the project page, and Tao’s launch post, retrieved 2026-07-26. The blueprint PDF carries a 2026-07-03 date, so the database is current. The launch paper has since been identified as arXiv 2501.16779, “New exponent pairs, zero density estimates, and zero additive energy estimates: a systematic approach” (Tao, Trudgian, and Yang, 2025-01-28), and is now cited; an earlier version of this entry recorded it as unidentified.

2 Terence Tao, launch post, 2025-01-28: “…abstracting out various relations between these exponents that were implicit in many papers in this subject, we were then able to run computer-assisted searches to improve some of the state of the art on these exponents in a largely automated fashion (without introducing any substantial new inputs from analytic number theory).”

Comparing efficiency rates across domains (2026)

This log’s synthesis. Not a source: a comparison assembled from the entries above, recorded here so the argument can cite it rather than rebuild it.

Left panel: horizontal bars showing the years each efficiency measurement covers, for seven physical technology cost curves and three AI algorithmic-progress estimates. Right panel: halving times on a log axis from 8 to 24 months, with 95% confidence intervals for the AI estimates and a dashed line at Moore's law.

Efficiency improvement rates, physical technologies against AI algorithmic progress.
  • How it was built. The physical rates are log-linear fits to each series in the Our World in Data CSV [→ OWID], computed by tools/sources_figures.py; a fit is used rather than an endpoint ratio so a single noisy observation cannot set the rate. The AI rates are quoted from their entries, which now sit in the algorithms section [→ Epoch on LMs, compute-to-AlexNet, ImageNet], not fitted here. Everything in this entry is therefore the log’s arithmetic over other people’s numbers.
  • AI algorithmic progress is fast, and not unprecedented. Language-model pretraining halves compute every 8 months and ImageNet every 9. DNA sequencing cost halved every 8.6 months over 2001–2013. The fastest measured AI efficiency rate and the fastest measured physical technology cost curve are the same number to within the confidence intervals.
  • Both are well inside a range that ordinary industrial technologies have occupied. Hard disk drives halved every 13 months, transistors every 17, DRAM every 19, laser diodes every 24. Compute-to-AlexNet at 16 months sits in the middle of that group rather than above it.
  • The distribution is heavily skewed, and the tail is where the interesting cases are. Of the 66 series, 65 fall in cost and one rises, but only five halve faster than every two years; the median falling series takes decades. Rapid exponential improvement is the exception across technologies even though it is the norm within the ones anybody writes about.
  • The measurement windows barely overlap, which limits what the comparison can support. The physical curves are mostly mid-twentieth-century and all end by 2013; the AI estimates all start in 2012 and none extends past 2023. Nothing here is a like-for-like contemporaneous comparison, and none of it covers the agent era at all.
  • What it is for. It fixes the baseline against which any AI-driven acceleration has to be judged. A theory that predicts AI bends the efficiency curve has to predict a bend relative to rates that were already this fast without AI. This is the log’s framing.
  • Dates: physical series 1929–2013; AI estimates published 2020-05-08, 2022-12-10, and 2024-03-09; figure generated 2026-07-26.
  • Bears on: Q1 growth rate, Q5 returns, Q7 incidence.
  • Links: generated by tools/sources_figures.py from posts/data/apple-picking/owid-66-technologies.csv (OWID source)
  • Unused: held for the introduction’s baseline argument; not yet cited.
  • Status: derived — every input figure is quoted in the entry it comes from, and the figure regenerates from the CSV and those quoted rates.

Epoch AI: ML hardware price-performance (2023)

Independent (research organization). The hardware curve the algorithmic-efficiency rates have to be set against. It is here rather than in the algorithms section because it is a comparator for all three domains, and because the comparison is the whole point: this is a slower curve than every algorithmic rate in the log.

  • The rate, with both samples. Epoch finds “computational price-performance [FLOP per $] doubling every 2.1 years for ML GPUs and 2.5 years for general GPUs.”
  • The samples. 47 ML hardware accelerators over 2010–2023, and 1,948 general GPUs over 2006–2021. Counts read from the page rather than quoted from a sentence.
  • Two caveats that bias the comparison in known directions. Measuring price-performance in FP32 may understate ML hardware, because ML workloads use lower-precision formats the metric ignores; and cluster hardware prices are often negotiated privately rather than listed, which makes accurate pricing difficult to determine. Both are Epoch’s points, given here in the log’s words because the fetched text could not be certified sentence-level verbatim — re-quote from the page before putting either inside quotation marks.
  • What it establishes for this log. A 2.1-year doubling is roughly 25 months, against 8 months for language-model pretraining efficiency, 9 for ImageNet, and 16 for compute-to-AlexNet [→ Epoch on LMs, ImageNet, compute-to-AlexNet]. So software has been the faster-moving term in AI for a decade, before any agent contributed to it. That matters for reading claims that AI will accelerate AI: the channel with the most historical headroom is the one AI agents are least demonstrably contributing to. This is the log’s reading.
  • Dates: published 2023-11-09; the ML accelerator series runs 2010–2023 and the general GPU series 2006–2021. Retrieved 2026-07-26.
  • Bears on: Q1 growth rate, Q5 returns, Q6 intertemporal.
  • Unused: not yet cited in the argument. It belongs alongside the rate comparison above, which currently compares AI algorithmic progress against physical technologies but not against the hardware it runs on.
  • Links: Epoch, trends in machine learning hardware
  • Status: verified — the headline rate quote and the sample counts read from the Epoch page, retrieved 2026-07-26. The two caveats are paraphrased from that page rather than certified verbatim, and are marked as such above.

Hao and co-authors: AI expands individual impact and contracts collective focus (2024)

Independent (academic). The largest bibliometric study of AI’s use in science: an LLM classifier labels AI-augmented papers across 41.3 million natural-science papers, and adopters are compared with non-adopters on output, citations, career timing, topic breadth, and follow-on engagement. It is correlational, and the effect sizes are large enough that the selection problem is the first thing to say about them.

  • The design, with the classifier’s own accuracy stated. “we used a pretrained language model to identify AI-augmented research, with an F1-score of 0.875 in validation against expert-labeled data. Using a dataset of 41.3 million research papers across natural science and covering distinct eras of AI, here we show an accelerated adoption of AI tools among scientists and consistent professional advantages associated with AI usage, but a collective narrowing of scientific focus.”
  • The individual-level associations, which are very large. “Scientists who engage in AI-augmented research publish 3.02 times more papers, receive 4.84 times more citations, and become research project leaders 1.37 years earlier than those who do not.”
  • The collective-level counterpart, in the opposite direction. “By contrast, AI adoption shrinks the collective volume of scientific topics studied by 4.63% and decreases scientist’s engagement with one another by 22.00%.”
  • The authors’ own reading, and it is close to this project’s thesis. “AI adoption in science presents a seeming paradox – an expansion of individual scientists’ impact but a contraction in collective science’s reach – as AI-augmented work moves collectively toward areas richest in data. With reduced follow-on engagement, AI tools appear to automate established fields rather than explore new ones, highlighting a tension between personal advancement and collective scientific progress.”
  • The identification problem, which the abstract does not address. No instrument or exogenous shock is claimed, so a threefold publication difference between adopters and non-adopters is equally consistent with productive scientists adopting AI first. The word in the abstract is “associated,” and it should be preserved in any use. The 4.63% and 22.00% figures are abstract-level and were not checked against the body. This caveat is the log’s.
  • Why it belongs here. “Moves toward areas richest in data” and “automate established fields rather than explore new ones” are the streetlight and depletion mechanisms this project is about, measured across all of natural science rather than in three domains. It is the widest-scope evidence in the log and the weakest-identified.
  • Dates: arXiv 2024-12-10, with versions through 2025-11-29; the arXiv listing notes acceptance at Nature, which this log has not confirmed with a DOI.
  • Bears on: Q1 growth rate, Q4 expertise, Q7 incidence.
  • Unused: not yet cited in the argument. It is the closest thing available to a field-wide test of the incidence prediction, and its correlational design is why the argument should cite it with the caveat attached rather than as a headline.
  • Links: arXiv 2412.07727
  • Status: verified-abstract — quotes and version history checked against the arXiv listing, retrieved 2026-07-26. Body figures not read; the design is a matched observational comparison, not an experiment.

Acemoglu: the simple macroeconomics of AI (2024)

Independent (academic). The standard low-end estimate of AI’s aggregate productivity effect, reached by a task-based application of Hulten’s theorem (Acemoglu 2024). It is in the log for two reasons: it is the number the growth debate is anchored on, and its own argument for why the estimate may be too high is an apple-picking argument in different words.

  • The headline, from the April 2024 version. “Using existing estimates on exposure to AI and productivity improvements at the task level, these macroeconomic effects appear nontrivial but modest—no more than a 0.71% increase in total factor productivity over 10 years.”
  • His reason for thinking even that is too high, which is the interesting part. “The paper then argues that even these estimates could be exaggerated, because early evidence is from easy-to-learn tasks, whereas some of the future effects will come from hard-to-learn tasks, where there are many context-dependent factors affecting decision-making and no objective outcome measures from which to learn successful performance. Consequently, predicted TFP gains over the next 10 years are even more modest and are predicted to be less than 0.55%.”
  • The exposure denominator. From the body: “0.23 × 0.199 = 4.6% of all tasks (or occupations) will be impacted by AI and computer” vision.
  • Three different headline figures are in circulation for this one paper, and the entry records all three. The April 2024 MIT version says 0.71% and less than 0.55%. The NBER working paper of May 2024 says “no more than a 0.66% increase in total factor productivity (TFP) over 10 years” and “less than 0.53%.” Aghion and Bunel characterize his result as about 0.07 percentage points a year [→ Aghion and Bunel]. Pin the version whenever the number is quoted; this log’s rule on revised figures is why all three are here rather than the latest.
  • The scope limit that matters most for this project. The model has no idea-production channel at all: AI enters only through task-level cost savings in the production of goods and services. So it cannot speak to AI’s contribution to research, which is what this log is about, and it should never be cited as though it could. This is the log’s point and it is the explicit motivation for the next entry.
  • Dates: MIT version 2024-04-05; NBER working paper 32487 issued May 2024; published in Economic Policy 40(121), 2025, which was not read here.
  • Bears on: Q1 growth rate, Q5 returns, Q7 incidence, Q8 benchmarks — the easy-to-learn-tasks argument is a claim about which benchmarks are informative.
  • Unused: not yet cited in the argument. The hard-to-learn-tasks passage is the most citable part, because it is an independent statement of the log’s own concern that measured gains come from the settings where measurement is easy.
  • Links: MIT PDF · NBER w32487
  • Status: verified — the MIT-version quotes read from that PDF and the NBER-version figures read from the NBER listing, retrieved 2026-07-26, so the discrepancy between them is checked rather than inferred. The published journal version was not read and may differ again.

Aghion and Bunel: a higher estimate, with the ideas channel explicitly omitted (2024)

Independent (policy note). The direct reply to Acemoglu, using a historical-analogy approach and a re-parameterized version of his own task-based formula (Aghion and Bunel 2024). It is the single most useful entry in the log for Q5, because the authors state in the abstract that the channel this whole project is about has been left out of both estimates.

  • Their estimates, and the omission, in one abstract. “Based on the first approach, we estimate that the AI revolution should increase aggregate productivity growth by between 0.8 and 1.3pp per year over the next decade. Using the second approach but with our own reading of the recent empirical literature on the various components of the task-based formula, we obtain a median estimate of 0.68pp additional annual total factor productivity (TFP) growth. Our estimates do not take into account the fact that AI automates tasks not only in the production of goods and services, our focus in this note, but also in the production of ideas.”
  • The spread of the second approach. Their reading of the same formula gives “aggregate productivity growth by between 0.07pp and 1.24pp, with a median estimate of 0.68pp,” against Acemoglu’s “much smaller extra growth estimate of 0.07 pp per year.”
  • They discount the Copilot trial the log also treats cautiously. “Peng et al. (2023) is less relevant because the task evaluated is too finely defined” [→ Copilot RCT]. So the disagreement with Acemoglu is partly about which micro estimates to feed the formula, not only about the formula.
  • Their passage on the ideas channel, quoted at length because it is the closest thing to this project’s thesis from growth economists. “AI could automate, or at least facilitate, the generation of new ideas (Aghion et al., 2018). It will thus help us generate new inventions and solve complex problems, as in the case of AlphaFold, which helps find new proteins, or GNoME, which suggests new materials that could be used in vehicles or everyday objects. The impact of AI on science and innovation is difficult to quantify, especially as AI’s ability to generate new ideas could face practical difficulties. For instance, it is not enough to identify several million potential new materials; they must still be validated experimentally. Nonetheless, AI will at the very least make the work of researchers easier. As AI tools gradually assist humans in identifying new hypotheses, designing protocols, and conducting experiments, the production of relevant ideas will increase. However, the time horizon of these effects remains highly uncertain.”
  • And the structural claim they attach to it. “These effects are leading to a permanent increase in the rate of productivity growth. The magnitude of this effect, however, is difficult to quantify.”
  • Why this entry justifies the log’s existence. Both sides of the headline macro debate omit the ideas channel and say so, and one of them names experimental validation as the binding difficulty — which is the validation-bottleneck thesis [→ validation bottleneck] appearing in a growth note two years before DeepMind’s essay. The quantity everyone agrees is missing is the quantity this log is trying to assemble. This is the log’s framing.
  • Dates: note dated June 2024, circulated via the Federal Reserve Bank of San Francisco.
  • Bears on: Q1 growth rate, Q5 returns, Q6 intertemporal, Q8 benchmarks.
  • Unused: not yet cited in the argument. It belongs in the scope section, as the reason a document about AI and research productivity is worth writing at all.
  • Links: PDF
  • Status: verified — all quotes read from the PDF, retrieved 2026-07-26. A policy note rather than a refereed paper, and the 0.8–1.3 point figure comes from a historical analogy to past general-purpose technologies rather than from estimation.

Toner-Rodgers, “AI, Scientific Discovery, and Product Innovation” (2024) — WITHDRAWN, do not cite

Disavowed (MIT). This November 2024 preprint reported that an AI materials-discovery tool deployed at a large R&D lab raised materials discovered by 44%, patent filings by 39%, and product innovations by 17%, with gains concentrated among the most able scientists. It became the most-cited empirical result on AI and scientific discovery, and specifically on how AI interacts with researcher expertise — which is exactly the question this project’s Q4 asks. It is logged here so that it is not re-used.

  • MIT disavowed the paper on 2025-05-16. Its Committee on Discipline stated it has “no confidence in the provenance, reliability or validity of the data” and “no confidence in the veracity of the research contained in the paper.” MIT asked that it be “withdrawn from public discourse.”
  • The author is no longer at MIT, and MIT asked arXiv and the QJE to withdraw it. Because arXiv accepts withdrawal requests only from authors, and the author had not submitted one, MIT wrote to arXiv directly.
  • Acemoglu and Autor, who had both publicly praised it, joined the disavowal. Their statement notes the paper was “already known and discussed extensively in the literature on AI and science, even though it has not been published in any refereed journal.”
  • Subsequent reporting indicates fabrication rather than error. The WSJ account describes an apparently invented study, including a faked data-use agreement, rather than a data-handling mistake.
  • Consequence for this project: Q4 has no headline finding. The most quotable claim about AI and researcher expertise is not evidence. Everything the argument says about expertise is therefore assembled from domain-specific observations, which is weaker but real.
  • Dates: preprint November 2024; MIT disavowal 2025-05-16; TechCrunch report 2025-05-17; retrieved 2026-07-26.
  • Bears on: Q4 expertise — as a negative entry, recording what cannot be used.
  • Links: MIT Economics statement · TechCrunch · WSJ account
  • Status: verified — the disavowal is verified against MIT’s own statement, retrieved 2026-07-26. The paper’s original figures are reproduced above only to make it identifiable and must not be cited as findings.

Conceptual models

Benjamin Jones: AI in R&D (2025)

Independent (economic model). Jones models research as a unit interval of complementary tasks whose outputs combine into progress. Humans can perform every task; machines can perform a share \(\gamma_t\), with machine productivity summarized by \(M_t\) and task bottlenecks governed by \(\theta\) (Jones 2025).

  • AI automates tasks, not necessarily whole results. When \(\gamma_t<1\), machines produce some task-level inputs while humans perform the remaining tasks. Strong complementarity means the human-only tasks can bottleneck the final research outcome even if AI becomes arbitrarily productive at its own tasks.
  • The human requirement is conditional, not absolute. In the model’s limiting case \(\gamma_t\to1\), machines take over all R&D tasks and progress no longer requires human research labor. Thus the model implies that partial task automation does not independently produce the composite result; it does not assume that human involvement is intrinsically necessary.
  • Breadth can matter more than intelligence. With bottlenecks, expanding the share of tasks AI can perform can accelerate progress more than extreme productivity gains on a narrow task subset.
  • No fixed frontier assumption. The framework explicitly studies increases in \(\gamma_t\) as AI takes over additional tasks. Its defining feature is task-level substitution plus complementarity, not a permanently fixed boundary between automatable and non-automatable tasks.
  • The idea production function is a CES over a unit measure of tasks. Equation (1) is \(\dot Z_t = \zeta Z_t^{\phi}\big[\int_0^1 r_t(j)^{\theta}dj\big]^{1/\theta}\) with \(\theta<0\), where \(\theta\) “governs the degree of complementarity between tasks - i.e., the strength of ‘bottlenecks’” and \(\phi\) decides whether “research advances might become easier (\(\phi > 0\)) or harder (\(\phi < 0\)) as progress continues.” Two notation traps for anyone reading this alongside the argument document: Jones’s \(\theta\) is the CES exponent (the argument calls that \(\rho\), and reserves \(\theta\) for the elasticity of substitution \(1/(1-\theta_{\text{Jones}})\)), and Jones separately uses \(\rho\) for the share of remaining tasks humans still perform.
  • Humans can perform every task, and within a task the two inputs are alternatives at a constant rate. Equation (2) sets \(r_t(j) = m_t(j)x_t(j)\) for \(0\le j<\gamma_t\) and \(r_t(j)=H\,l_t(j)\) for \(0\le j\le1\): “we imagine that humans can do all these tasks, but that machines have been created over time that perform some fraction of these tasks.” Machines are deployed only where cost-effective, \(m_t(j)/\mu_t \ge H/w_t\) (Assumption 1), the comparison being available at all because “research labor can do any task.”
  • Proposition 1 is a CES unit-cost function, which pins down the isoquant. Maximizing progress subject to \(D_t = \mu_t X_t^r + w_t L_t^r\) gives \(\dot Z_t/Z_t = \zeta Z_t^{\phi-1}D_t\big/\big[\gamma_t(\mu_t/M_t)^{\frac{\theta}{\theta-1}} + (1-\gamma_t)(w_t/H)^{\frac{\theta}{\theta-1}}\big]^{\frac{\theta-1}{\theta}}\), and “all heterogeneity in the machine-task productivities is summarized by the single index \(M_t\).” That denominator is the unit cost of the task aggregate, so the primal technology it is dual to is \(\big[\gamma_t y_1^{\theta} + (1-\gamma_t)y_2^{\theta}\big]^{1/\theta}\) over per-task outputs \(y_1 = (M_tX+HL_1)/\gamma_t\) and \(y_2 = HL_2/(1-\gamma_t)\). The duality is this log’s derivation; the paper states the cost side and does not draw the isoquant.
  • Infinite machine intelligence buys a bounded multiple, and that bound is the ceiling the argument tests. Corollary 3: “In the limit where \(M_t\to\infty\), the rate of progress increases by a multiple \(\eta_\infty = (1-s_t^X)^{-1/b}\) for \(\theta<0\),” with \(b=\theta/(\theta-1)\) and \(s^X_t\) the machine expenditure share. Jones’s illustration takes \(\theta=-1\) and \(s^X=1/3\), so the factor is \(2.25\): “the rate of progress would a bit more than double with infinite productivity across the entire current set of non-labor tasks.” Corollary 1 gives \(s^X_t = \gamma_t\) when machines and labor are exactly break-even per dollar (\(C_t = M_tw_t/H\mu_t = 1\)), in which case the bound reduces to \((1-\gamma_t)^{1/\theta-1}\) — computed here from the two results, not stated in the paper.
  • Dates: conference manuscript September 2025; NBER working paper 34312 issued 2025-10-02.
  • Bears on: Q2 autonomy, Q3 demand, Q5 returns, Q7 incidence.
  • Links: NBER working paper 34312 · September 2025 manuscript
  • Status: verified against the full manuscript.

Bazzichi, Riccaboni, and Castellacci: recombinant innovation (2026)

Independent (economic model). This April 2026 model puts ideas in a knowledge space and asks whether AI pushes R&D firms toward close, incremental recombinations or distant, radical ones.

  • Two AI margins point in different directions. Higher AI productivity can make distant combinations more feasible. But increasing the share of research tasks assigned to AI has an inverted-U effect: automation initially supports more radical combinations, then shifts research back toward incremental combinations once human–AI complementarity erodes.
  • Streetlight and duplication mechanisms. Heavy reliance on similar systems concentrates search in data-rich, well-explored regions (the “streetlight effect”) and makes different researchers converge on the same suggestions (the “stepping-on-toes effect”).
  • Strong limiting prediction. Under the paper’s functional form, optimal recombinant distance collapses to zero at full automation. This is not a generic theorem about AI; it follows from the model’s assumption that originality requires a residual complementary human contribution.
  • Distinct from apple-picking. Both can predict shallow and duplicate-heavy output, but this model locates the cause in homogenized search and lost human complementarity, not an intrinsic reach ceiling. It also predicts that greater AI productivity can increase radicalness even while broader automation eventually reduces it.
  • Dates: arXiv 2026-04-02.
  • Bears on: Q3 demand, Q5 returns, Q7 incidence.
  • Links: arXiv 2604.02189
  • Status: verified against the full paper.

DeepMind: conjecture machines and the validation bottleneck (2026)

Vendor (Google DeepMind policy essay, July 2026). A qualitative theory of AI-driven science: agents make hypotheses and candidate solutions abundant and cheap, while testing whether they survive contact with reality remains slow, costly, physical, and institutional.

  • Domain prediction. AI contributes end-to-end results first where checking is cheap and automatable — code execution, scoring functions, formal proof — while physical experiments, expert review, and tacit laboratory work become tighter bottlenecks elsewhere.
  • Candidate glut rather than researcher equivalence. More agents can generate more proposals without proportionally increasing validated knowledge. The relevant scarce input becomes verifier throughput, including peer review and experimental infrastructure.
  • Math is only partly exempt. Formalized proofs can be checked automatically, but natural-language proofs can still create “proof indigestion” faster than mathematicians can absorb them.
  • Relationship to Jones. This is a more specific bottleneck account, not a complete production model: it says which task should become scarce as generation gets cheap, but does not model the changing task boundary or aggregate growth.
  • Dates: published July 2026; the page shows no posting date, so this is the coarsest date in the log.
  • Bears on: Q3 demand, Q7 incidence.
  • Links: Conjecture Machines
  • Status: verified against the primary essay; vendor.

Davidson, Halperin, Houlden, and Korinek: recursive R&D feedback (2026)

Independent (economic growth model). A 2026 semi-endogenous growth model asks when automating AI research creates superexponential growth through an innovation network.

  • Two reinforcing loops. Better technologies raise research productivity in connected sectors (for example, software and hardware improve each other), while higher output finances more accumulable machine researchers. Together these loops can offset diminishing returns to ideas.
  • Bottlenecks matter, but need not dominate. Slow complementary sectors can block explosive growth; sufficiently fast expansion of task automation can relax those bottlenecks.
  • Different empirical object. The model predicts acceleration and cross-sector spillovers over time, not whether any one contribution is shallow, autonomous, duplicate, or easy to verify. The three domain snapshots in the companion argument therefore cannot adjudicate its central claim without a time series linking AI-generated improvements back into subsequent AI capability.
  • Dates: NBER working paper 35155 issued 2026-04-30.
  • Bears on: Q1 growth rate, Q6 intertemporal.
  • Links: NBER working paper 35155
  • Status: verified against the full working paper.

Aghion, Jones, and Jones: AI in the idea production function (2017)

Independent (economic model), peer-reviewed as a volume chapter. The foundational model of AI in knowledge production, and the direct ancestor of the ceiling test in the companion argument’s formalization (Aghion, Jones, and Jones 2019). AI enters the input index of the idea production function as the automated share of research tasks, and the elasticity of substitution does all the work. Anyone extending the argument’s formal section should start here rather than with the argument.

  • The organizing insight, which is Baumol’s. “One theme that emerges is based on Baumol’s ‘cost disease’ insight: growth may be constrained not by what we are good at but rather by what is essential and yet hard to improve.”
  • The functional form, in their notation. They write \(\dot A_t = A_t^{\phi}\big((B_tK_t)^{\rho} + (C_tS_t)^{\rho}\big)^{1/\rho} \equiv A_t^{\phi}F(B_tK_t, C_tS_t)\), “where \(S_t\) is the research labor used to make ideas,” with \(\beta_t\) the automated task fraction entering through \(B_t\) and \(C_t\). Their symbols are not the argument’s reserved notation: their \(\beta\) is the automated research-task share (the argument’s \(s\)), their \(\phi\) is the stock exponent, and their \(\gamma\) is an explosion index. Translate before citing.
  • Partial automation of research gives a level effect, not a growth effect — this is the ceiling result. “The second line follows if \(K_t/S_t\) is growing over time (i.e. if there is economic growth) and if the elasticity of substitution in \(F(\cdot)\) is less than one, which we’ve assumed. In that case, the CES function is bounded by its scarcest argument, in this case researchers. Automation then essentially produces a level effect but leaves the long-run growth rate of the economy unchanged if \(\phi<1\).”
  • Complete automation of idea production gives a singularity. “Once all tasks can be automated — i.e. once an A.I. replaces all people in the idea production function — the production of new ideas is given by \(\dot A_t = K_tA_t^{\phi}\). With \(\phi>0\), this differential equation is ‘more than linear.’… it is easy to see from this solution that \(A(t)\) exceeds any finite value before date \(t^* = 1/(\phi A_0^{\phi})\). This is a singularity.”
  • Whether full automation is required depends on the aggregator, and they are precise about it. “With the CES case and an elasticity of substitution less than one, we require that all tasks are automated. If only a fraction of the tasks are automated, then the scarce factor (labor) will dominate, and growth rates do not explode. We show in this section that with Cobb-Douglas production, a Type II singularity can occur as long as a sufficient fraction of the tasks are automated. In this sense, the singularity might not even require full automation.”
  • A pessimistic result of theirs that almost never gets cited. “In the Appendix we show that if some steps in the innovation process require human R&D, A.I. could possibly slow or even end growth by exacerbating business-stealing, which in turn discourages human investments in innovation.”
  • Why it is the most important conceptual entry in the log. The argument’s central formal claim — that task replacement’s ceiling is proportional to labor while apple-picking’s is not — is this paper’s bounded-by-the-scarcest-argument result, restated for a different stock. The argument reaches it independently and does not currently cite it. That is an omission rather than a disagreement. This assessment is the log’s.
  • Dates: NBER working paper 23928, October 2017; published as a chapter in The Economics of Artificial Intelligence: An Agenda, 2019.
  • Bears on: Q1 growth rate, Q2 autonomy, Q5 returns, Q6 intertemporal.
  • Unused: not yet cited in the argument, and it is the most consequential gap in the log. It belongs in the formalization section as the origin of the ceiling result, and in “Other theories” as the earliest entry.
  • Links: NBER w23928 · working paper PDF
  • Status: verified — every quote and equation above re-checked against the working-paper PDF’s own text, retrieved 2026-07-26. The Baumol sentence is from the abstract, which repeats it almost exactly in the introduction as “economic growth may be constrained not by what we do well but rather by what is essential and yet hard to improve”; the scarcest-argument, singularity, Cobb-Douglas, and business-stealing quotes are from the body and its footnote. The working paper is not peer-reviewed; the 2019 volume chapter is.

Korinek and Suh: whether wages collapse depends on the tail of task complexity (2024)

Independent (economic model). A transition model in which rising capability automates progressively more complex tasks (Korinek and Suh 2024). It is the closest published analogue to a reachable-set formulation, defined over a distribution of task complexity rather than over a stock of results — which makes it the natural comparison for apple-picking’s moving reach height.

  • Two regimes, turning on whether human task complexity is bounded. If the complexity distribution has a sufficiently thick infinite tail, “there is always enough work for humans, and wages may rise forever.” If human task complexity is bounded and full automation occurs, “wages collapse.”
  • The intermediate case is where the interesting behavior is. Automation productivity may generate “broad-based gains in the returns to all factors,” while “bottlenecks to growth from irreproducible scarce factors may exacerbate the decline in wages.”
  • How it relates to apple-picking. Both models make the reachable set a threshold on a fixed distribution and both let the threshold rise with capability. The difference is what the distribution is over: task complexity here, difficulty of results there. Korinek and Suh’s boundedness question — thick tail or bounded support — is exactly the question the argument’s alternative formalizations section asks about the difficulty density. The two literatures are asking one question in two vocabularies. This comparison is the log’s.
  • Dates: arXiv 2024-03-17; also circulated as an NBER working paper in 2024.
  • Bears on: Q2 autonomy, Q3 demand, Q5 returns, Q6 intertemporal.
  • Unused: not yet cited in the argument. It belongs in “Other theories” and in the alternative-formalizations discussion, where the thin-tail-versus-thick-tail question is currently raised without a citation.
  • Links: arXiv 2403.12107 · NBER listing
  • Status: verified-abstract — the quoted phrases and dates checked against the arXiv listing, retrieved 2026-07-26. The NBER working-paper number was seen in a search result rather than confirmed on nber.org, so confirm it before citing that number specifically.

Ide and Talamas: autonomous AI helps the most knowledgeable, assistive AI helps the least (2023)

Independent (economic model). A knowledge-hierarchy model in which AI agents can act as co-workers, solvers, or co-pilots (Ide and Talamas 2024). AI enters through which rung of the hierarchy it can staff, so capability determines the layer rather than scaling an input index. It is the best formal treatment in the log of the interaction between autonomy and expertise, which is the pair the argument’s Q4 row keeps running into.

  • The central result, and it is a conditional one. “We model AI as a technology that converts computational resources into ‘AI agents’ that operate autonomously (as co-workers and solvers/co-pilots) or non-autonomously (solely as co-pilots). Autonomous AI primarily benefits the most knowledgeable individuals; non-autonomous AI benefits the least knowledgeable. However, output is higher with autonomous AI. These findings reconcile contradictory empirical evidence and reveal tradeoffs when regulating AI autonomy.”
  • Why the reconciliation claim matters for this log. The expertise section holds studies that disagree about sign — compression in customer support and in an education-gap experiment, divergence among Kenyan entrepreneurs [→ support agents, education gap, Kenya]. This model says the sign should depend on whether the AI acts autonomously or assists, which is a structured prediction rather than an appeal to setting. Nothing in the log tests it, and it is testable. This reading is the log’s.
  • Dates: arXiv 2023-12-09, with twelve versions through 2025-05-17; a journal DOI is listed on the arXiv page.
  • Bears on: Q2 autonomy, Q3 demand, Q4 expertise.
  • Unused: not yet cited in the argument. It belongs in “Other theories,” and it is the one entry there that would give the Q4 row a mechanism rather than a list of conflicting results.
  • Links: arXiv 2312.05481 · journal DOI
  • Status: verified-abstract — the abstract, authorship, and version history checked against the arXiv listing, retrieved 2026-07-26. Twelve versions over eighteen months, so any result quoted from the body needs its version pinned.

Besiroglu, Emery-Xu, and Thompson: AI-augmented R&D is more capital-intensive (2022)

Independent (academic, peer-reviewed in Research Policy). Estimates idea production functions for two computer-vision tasks and reads off factor shares (Besiroglu, Emery-Xu, and Thompson 2024). It is the closest thing in the log to a measured elasticity of research output with respect to compute, which is the parameter the argument’s Q5 row is about.

  • The finding, with the authors’ conclusion stated as conditional. “We assess this impact by estimating the idea production function for AI in two computer vision tasks that are considered key test-beds for deep learning and show that AI idea production is notably more capital-intensive than traditional R&D. Because increasing the capital-intensity of R&D accelerates the investments that make scientists and engineers more productive, our work suggests that AI-augmented R&D has the potential to speed up technological change and economic growth.”
  • What this log will not claim from it. The frequently-repeated summary that US growth might double, and the specific estimated capital shares, are not in the abstract and have not been checked against the body here. Do not attach a doubling figure to this entry.
  • Why it bears on the argument’s formal section. Every one of the four theories is a statement about how AI spend enters the input index, and the capital share of idea production is the empirical content of that statement. Two computer-vision tasks is a thin basis for it, but it is more than the argument currently has, which is nothing. This framing is the log’s.
  • Dates: arXiv 2022-12-15 (v2 2023-01-02); published in Research Policy 53(7), 2024.
  • Bears on: Q1 growth rate, Q5 returns, Q6 intertemporal.
  • Unused: not yet cited in the argument. It is the natural empirical anchor for the exchange-rate parameter the formalization introduces and never calibrates.
  • Links: arXiv 2212.08198
  • Status: verified-abstract — abstract, dates, and journal details checked against the arXiv listing, retrieved 2026-07-26. The publisher blocks automated fetching, so the body and its estimates have not been read.

Cyber

UK AISI: multi-step cyber attack scenarios (2026)

Government (UK AI Security Institute). Evaluates agents on two purpose-built cyber ranges, varying model generation and inference-time token budget (Folkerts et al. 2026).

Correction. This log previously filed the paper under METR. It is AISI’s: every author carries the affiliation “AI Security Institute, United Kingdom,” and METR appears in the paper only in related work, as Kinniment et al. (2024). One quotation was also wrong, and is corrected in the ranges bullet below.

Scatter plot of average steps completed against model release date from GPT-4o in August 2024 to Opus 4.6 in February 2026, rising from under 2 to about 15.6, with horizontal grey lines marking nine attack milestones.

Average steps completed on the 32-step corporate range by model release date, at 10M and 100M token budgets.

The paper’s Figure 2. Filled markers are the 100M-token budget and open markers the 10M budget; the grey horizontal lines mark the nine milestones of the attack chain. Over eighteen months the 100M-token average rises from under 2 steps to about 15.6, which falls between milestones 4 and 5 of 9. The vertical distance between a model’s two markers is the effect of a tenfold compute increase at fixed capability.

Grid of per-step completion rates for seven models; each row is a different starting point in the 32-step chain. Cells are dark at low step numbers in the start rows, and mid-range start rows show completions at later steps that are zero in the start rows.

Per-step completion rates at a 10M-token budget for agents started at the beginning of the range and for agents started after each milestone.

The paper’s Figure 4. Each row is a starting condition — either the beginning of the range or immediately after a named milestone — with model, scaffold, and the 10M-token budget held constant; each cell gives the fraction of runs completing that step. Comparing rows within a model varies where the agent starts while holding the target fixed.

  • Log-linear compute scaling, no plateau. “Model performance scales log-linearly with inference-time compute, with no observed plateau — increasing from 10M to 100M tokens yields gains of up to 59%, requiring no specific technical sophistication from the operator.” The final clause is the one with policy content: the gains do not require an expert operator.
  • The ranges, and why they were built this way. “Two purpose-built cyber ranges—a 32-step corporate network attack and a 7-step industrial control system attack—that require chaining diverse capabilities over long attack sequences.” Chaining is the point; single-step capability is not what is being measured. Correction: this log previously gave the second half as “chaining heterogeneous capabilities across extended action sequences,” which is not the authors’ wording; the sentence above is the abstract’s.
  • The best run, in human-time units. “The best single run completed 22 of 32 steps, corresponding to roughly 6 of the estimated 14 hours a human expert would need.” This is the only place in the log where a cyber result is denominated in expert-hours, which makes it the natural bridge to the METR horizon series [→ METR horizons].
  • The second range is close to a floor, and that asymmetry is informative. “On the industrial control system range, performance remains limited, though the most recent models are the first to reliably complete steps, averaging 1.2–1.4 of 7 (max 3).” Same models, same budgets, two domains an order of magnitude apart in yield.
  • Per-generation gains. On the 32-step corporate range, average steps completed at a 10M-token budget rose from 1.7 (GPT-4o, Aug 2024) to 9.8 (Opus 4.6, Feb 2026); “each successive model generation outperforms its predecessor at fixed token budgets.”
  • Agents can complete late steps they cannot reach. The starting-point experiment is the paper’s own summary of Figure 4: “Mid-range starts enable completion of steps that models fail to reach in end-to-end runs, suggesting that models struggle with long-horizon tasks even when capable of solving constituent steps in isolation.” This matters for how a ceiling should be read here. Part of what looks like an unreachable band is a chaining limit rather than a difficulty limit, and chaining is a property of the task’s length, not of the result’s depth.
  • But difficulty is also intrinsic, not only positional. “Performance drops on later steps regardless of starting point.” The authors name two mechanisms, context accumulation and inherent step difficulty, and quantify the second: “Milestones 1–2 typically require 2–4 actions each; later steps demand sequences of 8–16 actions and domain expertise in areas such as Active Directory exploitation or CI/CD pipeline manipulation.” So both effects are present, and this single experiment does not apportion them.
  • Dates: arXiv 2026-03-11 (v3 2026-03-17); the model range runs GPT-4o (August 2024) to Opus 4.6 (February 2026).
  • Bears on: Q1 growth rate, Q2 autonomy, Q5 returns, Q6 intertemporal, Q7 incidence.
  • Links: arXiv 2603.11214
  • Status: verified against the v3 full text; Figures 2 and 4 reproduced from the paper.

Google Big Sleep (2025)

Vendor (Google). Autonomous vulnerability-discovery agent from Google DeepMind and Project Zero.

  • Live SQLite zero-day (July 2025). Found CVE-2025-6965, a memory-corruption flaw affecting all SQLite versions before 3.50.2, that was “known only to threat actors and… at risk of being exploited,” and cut it off before exploitation. Google framed it as “the first time an AI agent has been used to directly foil efforts to exploit a vulnerability in the wild.”
  • 20 novel open-source bugs (Aug 4 2025). Twenty previously-unknown vulnerabilities in widely-used OSS (FFmpeg, ImageMagick, etc.), “found and reproduced autonomously,” with a human only in final review.
  • Caveat. Big Sleep’s original Nov-2024 SQLite find was in a development branch before release, so it never hit production — a genuine unknown bug but not an in-the-wild zero-day. The July-2025 CVE is the stronger “real threat” case.
  • Dates: first SQLite find November 2024, in a development branch; CVE-2025-6965 announced 2025-07-15; the twenty-bug disclosure 2025-08-04.
  • Bears on: Q2 autonomy, Q4 expertise.
  • Links: cloud.google.com · blog.google · TechCrunch
  • Status: quotes verified by direct fetch; figures are vendor self-reports.

DARPA AIxCC finals (2025)

Government (DARPA). AI Cyber Challenge finals, DEF CON, August 2025 — the strongest non-commercial anchor for autonomous discovery, and the only source in this log that measures the same task set at two dates a year apart.

Grouped bar chart comparing August 2024 semifinal and August 2025 final rates for synthetic bugs identified and identified bugs patched.

Semifinal and final performance on the same competition structure, twelve months apart.

Derived from the figures in this entry rather than reproduced from DARPA. Both rates roughly doubled or better across twelve months on a fixed competition structure — but the teams rebuilt their systems between the two, so this is progress in the whole stack rather than in the models alone.

  • Seven fully autonomous systems, 54 million lines of code. Team Atlanta won, Trail of Bits second, Theori third. Systems ran without human intervention on real open-source software.
  • 18 genuine zero-days, with a denominator alongside. Teams found 18 real, non-synthetic vulnerabilities — 6 in C codebases (one of which maintainers found and patched in parallel) and 12 in Java — and supplied 11 patches for them. Separately, on the competition’s synthetic bugs they found 54 of the 63 planted across the challenges and patched 43.
  • A dated year-over-year capability comparison on a fixed benchmark. Against the August 2024 semifinal, synthetic-vulnerability identification rose from 37% to 86% and the patch rate among identified bugs from 25% to 68%. In the semifinal, teams were markedly better on C than Java; by the final the success rates across the two had converged. This is the cleanest twelve-month capability series in the log, because the organizers held the competition structure fixed.
  • Cost per task is reported, which almost nothing else here does. About $152 per competition task, against bug bounties that “can range from hundreds to hundreds of thousands of dollars.” Each team received $50,000 in model credits from each of Anthropic, Google, and OpenAI for the final.
  • Speed. Patches were submitted in an average of 45 minutes.
  • DARPA revised its own headline after publication. The original write-up said the final contained 70 synthetic vulnerabilities; the competition administrator later determined it was 63. The count discovered (54) did not change, so the rate moved from 77% to 86%. Recorded here per this log’s rule on revised figures — the pre-correction number is still in circulation.
  • Dates: semifinal competition August 2024; final competition and results announcement 2025-08-08; figures revised after publication, see the correction bullet.
  • Bears on: Q1 growth rate, Q2 autonomy, Q5 returns, Q6 intertemporal, Q8 benchmarks. Previously tagged Q2 only; the year-over-year comparison and the per-task cost widen it.
  • Links: DARPA results release · AIxCC site
  • Status: verified against the DARPA release including its post-publication correction note, retrieved 2026-07-26. Figures are the organizer’s own; no independent audit of the scoring exists.

CyberGym (2025)

Independent (academic benchmark). 1,507 historical vulnerabilities across 188 projects; agents attempt to reproduce them.

  • A ceiling, but the recorded number was stale. This entry previously gave an 11.9% reproduction ceiling from the version as first read. The current v3 abstract says “even the top-performing combinations only achieve a ~20% success rate, demonstrating the overall difficulty of CyberGym.” Use ~20%; treat 11.9% as superseded rather than contradicted, since both the benchmark and the frontier models moved between versions.
  • Incidental zero-days, also revised upward. The earlier reading recorded 15. v3: “we show that CyberGym leads to the discovery of 34 zero-day vulnerabilities and 18 historically incomplete patches.” The incomplete patches were missing from this entry altogether and are the more interesting half — cases where a human fix did not actually close the hole.
  • What the task actually is, which limits the comparisons it supports. CyberGym “primarily tasks agents with generating a proof-of-concept test that reproduces a vulnerability, given only its text description and the corresponding codebase.” That is reproduction from a description, not discovery from scratch, so the success rate is not commensurable with Big Sleep or Mythos figures [→ Big Sleep, Mythos].
  • Scale. “A large-scale benchmark featuring 1,507 real-world vulnerabilities across 188 software projects.”
  • The authors’ stated motivation. Existing evaluations “fall short, because they are based on small-scale benchmarks and only measure static outcomes, failing to capture the full, dynamic range of real-world security challenges.”
  • Dates: arXiv 2025-06-03, current version v3 2026-03-24. Figures above are from v3, re-read 2026-07-26.
  • Bears on: Q2 autonomy, Q5 returns, Q7 incidence.
  • Links: arXiv 2506.02548
  • Status: verified-abstract.

Anthropic Mythos preview (2026)

Vendor (Anthropic). Frontier cyber model preview (April 2026); all figures are self-reports.

  • Headline claims. Mythos “autonomously discovered thousands of previously unknown vulnerabilities,” including a 27-year-old OpenBSD DoS and a 16-year-old FFmpeg out-of-bounds write that fuzzers had exercised “5 million times without triggering.”
  • Cost accounting (vendor-reported). OpenBSD 27-year bug: found across “~1,000 scaffold runs at a total cost under $20,000”; the single successful run “cost under $50” — Anthropic stresses $50 “only makes sense with full hindsight, since… they can’t know in advance which run will succeed.” FFmpeg: several vulnerabilities after “several hundred runs… roughly ten thousand dollars”; three fixed in FFmpeg 8.1. N-day exploitation: “developing a full root exploit from a known vulnerability costs under $1,000 and takes half a day.”
  • Depth claim. Anthropic presents the OpenBSD SACK bug as requiring reasoning about signed integer overflow — i.e. not a one-step mechanical find.
  • Dates: preview posted early April 2026. The page metadata reads 2026-04-09, but AISLE’s response to it is dated 2026-04-07, so the metadata is probably a modification date and the true posting date is unresolved.
  • Bears on: Q1 growth rate, Q7 incidence, Q8 benchmarks.
  • Links: red.anthropic.com
  • Status: unverified (quotes not yet checked against the primary post); vendor. The “thousands” figure is unaudited.

XBOW on HackerOne (2025)

Vendor (XBOW), with public leaderboard data. Autonomous pentester that topped a HackerOne leaderboard.

  • Steep sub-linearity in value per run. Of ~1,060 submissions, only ~132 were confirmed and resolved, with ~208 duplicates and ~209 “informative.” Re-running the picker mostly re-finds picked apples.
  • Composition. Finds are dominated by web-app classes (RCE, info disclosure, cache poisoning, SQLi) on a large, heterogeneous, under-audited attack surface.
  • Dates: leaderboard placement reported June 2025; the skeptical companion piece is dated 2025-06-29; XBOW’s own post carries no visible date.
  • Bears on: Q3 demand, Q5 returns, Q7 incidence.
  • Links: xbow.com · skeptical: raw.pm
  • Status: unverified; vendor. The submission and duplicate counts are XBOW’s own and have not been checked against the HackerOne leaderboard here; the skeptical companion piece disputes the framing rather than the counts. An earlier version of this line gave the provenance without saying how far the figures had been checked, which left the entry unclassifiable.

curl ends its bug bounty (2026)

Independent (maintainer report). curl ended its bug-bounty program (Jan 2026) after “AI slop” swamped triage.

  • AI-assisted submissions reached ~20% of entries in 2025 while genuine-vulnerability yield fell; Daniel Stenberg likened the flood to a DoS. Evidence that validation/triage, not discovery, is becoming the binding constraint.
  • Dates: program closed January 2026, reported 2026-01-21; the ~20% AI-assisted share refers to 2025 submissions.
  • Bears on: Q3 demand, Q5 returns.
  • Links: The Register
  • Status: verified (secondary reporting of a first-party announcement).

Gemini and developer experience: security of the resulting code (2026)

Independent (academic, controlled study). A quantitative programming study with 159 developers recruited on Upwork, assigned a security-related task with no AI, free Gemini, or paid Gemini. It is the only entry in this log that measures AI’s effect on a security outcome by user experience, which is the Q4 question in the cyber domain.

  • AI access did not significantly improve code security. No significant difference between the AI conditions and the control on the security of the resulting software, and none between the free and paid versions.
  • Prior experience did, and was not substitutable. Programming experience significantly improved code security, and the authors conclude it “cannot be fully substituted by Gemini.” Developers with no security experience did show improvement when supported by Gemini, so the effect is a partial floor-raise rather than a ceiling-raise.
  • The direction contrasts with the discovery-side results. Everything else in the cyber section measures finding vulnerabilities, where AI performs well; this measures not introducing them, where it does not. If the same models are strong at detection and neutral at prevention, the net effect on the stock of bugs in the world is not signed by the discovery evidence alone.
  • Scope caveats. Freelancer recruitment, one task, one tool, and a null result on a sample of 159 — underpowered for small effects, and the absence of a difference is not evidence of equivalence.
  • Dates: arXiv 2026-03-16 (v2 2026-03-17).
  • Bears on: Q3 demand, Q4 expertise.
  • Unused: not yet cited in the argument. It is the best available answer to “how expert are the people who do cyber work with AI” and would replace an inference currently drawn from who reviews and who scaffolds.
  • Links: arXiv 2603.15298
  • Status: verified-abstract — checked against the arXiv abstract and conclusion, retrieved 2026-07-26.

Cyber labour-market indicators (2026)

Industry surveys and job-posting data; not academic. Two 2026 datasets on what AI is doing to demand for security staff. Neither is a causal estimate and both come from parties with an interest in the answer, but together they are the only quantitative purchase this log has on Q3 in the cyber domain, where the argument currently reasons from curl’s bug bounty and ARTEMIS’s hourly rates.

  • Headcount reduction is rare; role composition is shifting. The SANS 2026 workforce survey of nearly 1,000 practitioners across six regions reports 49% of organizations seeing reduced manual analysis time and 48% workflow automation gains, but only 16% reporting actual headcount reduction. Among organizations reporting role changes, reductions concentrate in SOC and security analysts (32%), threat intelligence analysts (26%), and incident responders (22%).
  • The reductions fall on the entry level, and new senior categories are appearing. Among organizations adding roles, 34% added AI/ML security specialists, 32% AI security engineers, and 30% AI governance analysts. SANS reports expert-level roles are now the hardest to fill. The stated concern is that AI is automating precisely the junior work through which the next generation trained.
  • Job postings show the shift as early and small. An analysis of 5,197 AppSec postings across 796 companies from October 2025 to May 2026 finds AI mentioned in 20.8% of enriched descriptions, with prevalence rising from 2.1% in November 2025 to 7.2% in May 2026, and a 10.9% salary premium on AI-mentioning roles. Only 1.6% of all postings carry AI terms in the title.
  • How to read this against the theories. Composition shifting toward senior judgement while entry-level enumeration contracts is task replacement’s reallocation prediction, not human replacement’s level effect. It is also consistent with apple-picking, where the reachable zone is exactly the routine work. It does not discriminate between them.
  • Treat both as weak evidence. A self-selected practitioner survey and a job-posting scrape measure stated intentions and advertising language, not employment or wages. Neither has a counterfactual.
  • Dates: SANS 2026 Cybersecurity Workforce Research Report presented at RSAC 2026, with secondary coverage 2026-07-22; its exact release date is not established. The AppSec posting analysis is dated 2026-05-26 and covers October 2025 to May 2026.
  • Bears on: Q3 demand, Q4 expertise.
  • Unused: not yet cited in the argument. It is the closest thing to the occupation-level evidence the Known gaps section asks for, and should be used there with its weakness stated rather than left out.
  • Links: SANS announcement · Help Net Security coverage · AppSec hiring analysis
  • Status: unverified — figures are quoted from the publishers’ own summary pages, retrieved 2026-07-26; neither underlying dataset has been inspected and neither is peer-reviewed.

Vidoc Security reproduction (2026)

Independent (security firm). Reproduced the finding of Mythos-class bugs with public models “but they didn’t build the weapon” — detection is cheap, deep exploit development is hard.

  • Dates: posted 2026-04-09 per the page’s own metadata, within about two days of the Mythos preview — which bounds how much reproduction work the claim can represent.
  • Bears on: Q1 growth rate, Q4 expertise.
  • Links: vidocsecurity.com
  • Status: unverified quote.

AISLE: the jagged frontier (2026)

Independent (security research firm). Post-Mythos analysis of model capability across isolated cyber tasks.

  • Rankings reshuffle. “There is no stable best model across cybersecurity tasks… capability rankings reshuffle completely across different security tasks.” GPT-OSS-120b recovered the public OpenBSD SACK chain but failed a simple Java ArrayList data-flow trace, while several smaller models succeeded.
  • Cheap detection. In AISLE’s isolated-code test, even a model with 3.6B active parameters priced at $0.11 per million tokens detected the showcased buffer overflow — detection on preselected code, not repo-scale discovery.
  • Caveat. Tests used isolated code and plain API calls rather than end-to-end autonomous discovery.
  • Dates: posted 2026-04-07.
  • Bears on: Q7 incidence, Q8 benchmarks.
  • Links: aisle.com
  • Status: verified against the primary post.

Palo Alto Networks Mythos deployment (2026)

Vendor claim, contested. The Information reported Mythos at Palo Alto Networks “found more than two dozen critical vulnerabilities in about three weeks… burned through more than $1 million worth of tokens.”

  • Implied costs. The offcuts post plots this at ≈$42k/bug average, with marginal cost approaching a guessed human rate of ≈$100k/bug.
  • Skeptical note. The “5×” / “75 bugs vs 5–10/month” framing is contested as a reporting/attribution artifact (26 CVEs counted as 75 issues). Treat vendor “N× more bugs” framing as marketing until independently confirmed.
  • Dates: deployment described as about three weeks, spring 2026; Axios follow-up 2026-05-13. The Information’s original is paywalled and its date is not confirmed here.
  • Bears on: Q1 growth rate, Q8 benchmarks.
  • Links: The Information · flyingpenguin.com · Axios
  • Status: unverified; numerator, denominator, and token spend not reproducible from primary inputs.

UK AISI Mythos runs, via dbreunig (2026)

Independent runs, secondary summary. Drew Breunig’s summary of the UK AI Security Institute’s Mythos evaluations.

  • No diminishing returns at 100M tokens. “None of the models given a 100M budget showed signs of diminishing returns” — the basis for his “cybersecurity is proof-of-work now” framing (you win by spending more).
  • Budgeting. “100M tokens per attempt, $12,500 per Mythos attempt, $125k for all ten runs.”
  • Dates: Breunig’s summary posted 2026-04-14; the underlying AISI runs are not separately dated on the page.
  • Bears on: Q4 expertise, Q5 returns, Q8 benchmarks.
  • Links: dbreunig.com
  • Status: verified against the blog post; underlying AISI data not independently checked.

Lyptus: offensive-cyber time horizons (2026)

Independent (research org). Estimates human-equivalent time horizons for offensive cyber tasks, METR-style.

Log-scale plot of P50 offensive-cyber time horizon against model release date from GPT-2 in 2019 to Opus 4.6 in 2026, with a fitted trend and confidence band.

P50 time horizon for offensive-cyber tasks against model release date.

Note the methodological inheritance: this is METR’s construction applied to cyber tasks, so it and the general horizon series are not independent confirmations of each other. The confidence band at the right edge is wide enough to accommodate most claims about acceleration.

  • Horizon growth. From ~30 s (GPT-2, 2019) to ~3 h (Opus 4.6 / GPT-5.3 Codex, 2026), doubling every ~9.8 months and accelerating to ~5.7 months since 2024.
  • Point estimates. P50 horizons of 3.1 h for GPT-5.3 Codex and 3.2 h for Opus 4.6 at a 2M-token budget. Re-running GPT-5.3 Codex failures at 10M tokens raised its estimated P50 to 10.5 h, but with a very wide 2.4–63.5 h interval because the task set was near saturation.
  • Saturation. GPT-5.5 saturated the task set, making further horizon estimates impossible on that benchmark — a measurement ceiling, not evidence that capability plateaued.
  • Dates: main study 2026-04-02; saturation update 2026-05-27; the horizon series runs from GPT-2 (2019) to Opus 4.6 and GPT-5.3 Codex (2026).
  • Bears on: Q6 intertemporal, Q8 benchmarks.
  • Links: lyptusresearch.org · saturation update
  • Status: verified against the primary pages.

ARTEMIS pentest study (2025)

Independent (academic). Head-to-head pentest comparison of AI variants and human professionals.

Grid with vulnerabilities as rows and twelve discoverers as columns; dark cells mark discoveries, scattered unevenly across both rows and columns.

Which of thirteen vulnerabilities each of ten human participants (P1-P10) and two ARTEMIS variants (A1, A2) discovered.

The paper’s Figure 4. Rows are the thirteen vulnerabilities found by anyone, columns are the ten human participants and the two ARTEMIS variants, and a filled cell means that discoverer found it. The paper gives no prose reading of the figure, so the following is read off the chart rather than quoted: one vulnerability (idrac-default-creds-2) was found only by the two agents; the one found by the most humans (tinypilot-windows-rce, seven of ten) was found by neither agent; and every other agent find is shared with at least one human.

  • The comparison, stated exactly. “We evaluate ten cybersecurity professionals alongside six existing AI agents and ARTEMIS, our new agent scaffold, on a large university network consisting of ~8,000 hosts across 12 subnets.” The authors describe it as “the first comprehensive evaluation of AI agents against human cybersecurity professionals in a live enterprise environment.”
  • The result. “ARTEMIS placed second overall, discovering 9 valid vulnerabilities with an 82% valid submission rate and outperforming 9 of 10 human participants.” The valid-submission rate matters as much as the rank, because it is the false-positive discipline that usually separates agent output from professional output.
  • The scaffold does the work, not the model class. “While existing scaffolds such as Codex and CyAgent underperformed relative to most human participants, ARTEMIS demonstrated technical sophistication and submission quality comparable to the strongest participants.” Six other agents were on the same network and lost; only the purpose-built one won. Any reading of this entry as “AI beats pentesters” has to survive that sentence.
  • The cost gap. “AI agents offer advantages in systematic enumeration, parallel exploitation, and cost — certain ARTEMIS variants cost $18/hour versus $60/hour for professional penetration testers.” Note “certain variants”: this is the cheap end of their own range, not an average.
  • The authors’ stated capability gaps. “AI agents exhibit higher false-positive rates and struggle with GUI-based tasks.”
  • Dates: arXiv 2025-12-10 (v2 2026-03-03).
  • Bears on: Q3 demand, Q8 benchmarks.
  • Links: arXiv 2512.09882
  • Status: verified-abstract. Figure reproduced from the source.

DARPA Cyber Grand Challenge: the pre-LLM autonomy baseline (2016)

Government (DARPA). Fully autonomous find-and-patch, demonstrated nine years before the AIxCC final and with no language model anywhere in it. It is the control every 2025–2026 autonomy claim needs: machines were already doing unaided discovery and repair on unseen software using symbolic execution and fuzzing, so what changed since is the breadth and cost of the technique, not the existence of autonomy.

  • The first all-machine tournament, and the organizers’ own account of what surprised them. DARPA describes it as “the world’s first all-machine cyber hacking tournament.” Mike Walker, the program manager, on the result: “I am amazed at the speed with which the machines responded to the use of bugs in software they had never seen before and fielded patches in response.”
  • Placings and prizes. Mayhem (ForAllSecure) first at $2 million, Xandra (TECHx) second at $1 million, Mechanical Phish (Shellphish, UC Santa Barbara) third at $750,000.
  • Why it belongs in this log rather than in a history section. The AIxCC final reports an 86% synthetic-vulnerability identification rate in 2025 [→ AIxCC] against a 2016 predecessor that already ran unaided; treating 2025 as the origin of autonomous vulnerability discovery overstates the change by nine years. What a like-for-like comparison would need is the two competitions’ task sets put on one scale, which nobody has done. Both points are this log’s, not DARPA’s.
  • Dates: final event 2016-08-04; DARPA announcement 2016-08-05, updated 2016-08-07.
  • Bears on: Q1 growth rate, Q2 autonomy, Q4 expertise — the expertise here is classical program analysis rather than model capability.
  • Unused: not yet cited in the argument. It belongs wherever the argument treats autonomous discovery as new, and it is the natural companion to the AIxCC entry.
  • Links: DARPA results release
  • Status: verified — quotes and prize amounts checked against the DARPA release, retrieved 2026-07-26.

Google Project Zero, “Project Naptime”: tooling versus the model (2024)

Independent (Google Project Zero). The predecessor harness to Big Sleep [→ Big Sleep], and the cleanest isolated measurement in the log of how much of an apparent capability gain comes from giving the model tools rather than from the model. Same benchmark, same models, scaffold varied.

  • Scaffolding alone moved the benchmark by up to a factor of twenty. The team reports it “increased CyberSecEval2 benchmark performance by up to 20x from the original paper.” On the Buffer Overflow category, GPT-4 Turbo pass@20 rose from a reproduced “0.20” to “1.00”; on Advanced Memory Corruption, from “0.42” to “0.76”.
  • The authors’ own statement of what the benchmark leaves out, which is the reason this is a Q8 entry. They describe the tasks as “closer to the typical usage of targeted, domain-specific fuzzing performed as part of a manual review workflow than a fully autonomous researcher,” because “a large part of security research is finding the right places to look” — and the benchmark hands the model the place to look.
  • Why it matters for reading everything else in this section. A twentyfold benchmark swing from harness changes at fixed model capability means dated capability series that vary scaffold and generation together cannot attribute their slope to either. That is the confound the log’s Q6 gap describes, measured here directly. This reading is the log’s.
  • Dates: posted 2024-06-20.
  • Bears on: Q1 growth rate, Q2 autonomy, Q6 intertemporal, Q8 benchmarks.
  • Unused: not yet cited in the argument. It is the strongest available evidence for the argument’s claim that harness effects can swamp generation effects, currently supported only by PERFOPT and TTT-Discover in the algorithms domain.
  • Links: Project Zero
  • Status: verified — figures and both quotes checked against the primary post, retrieved 2026-07-26.

Fang and co-authors: autonomous exploitation of one-day vulnerabilities (2024)

Independent (academic, University of Illinois Urbana-Champaign). Tests whether an agent can write and run a working exploit given only a public vulnerability description. The headline is high and the entry is here mostly for the sentence that follows it, which is the sharpest published statement of how much a human-written description contributes to an “autonomous” result.

  • 87% with the CVE description, 7% without. “GPT-4 is capable of exploiting 87% of these vulnerabilities compared to 0% for every other model” — and the same abstract reports that “without the description, GPT-4 can exploit only 7% of the vulnerabilities.” The authors read the gap as reassuring rather than as a limitation: “Fortunately, our GPT-4 agent requires the CVE description for high performance.”
  • The denominator is small enough to matter. “we collected a dataset of 15 one-day vulnerabilities,” so 87% is thirteen of fifteen and a single case moves the figure by about seven points.
  • What the 87%/7% split measures. The human-written CVE description contains the localization that the Naptime post identifies as most of security research [→ Naptime]. A twelvefold drop when it is withheld is the closest thing in the log to a decomposition of an autonomy claim into the model’s contribution and the human substrate’s. This framing is the log’s.
  • Dates: arXiv 2024-04-11 (v2 2024-04-17); the models tested are the early-2024 frontier, which is well behind.
  • Bears on: Q2 autonomy, Q4 expertise, Q7 incidence.
  • Unused: not yet cited in the argument. It is the best single citation for the argument’s closing caveat that “autonomous” almost always means autonomous execution after humans chose the target.
  • Links: arXiv 2404.08144
  • Status: verified-abstract — all four figures and the quotes checked against the arXiv abstract, retrieved 2026-07-26. The body has not been read and the result has not been independently replicated.

Fang and co-authors: teams of agents on zero-day vulnerabilities (2024)

Independent (academic, same group as the one-day paper). The follow-up, which drops the CVE description and adds orchestration: a planner agent dispatching specialized subagents.

  • The problem it targets, in the authors’ words. Single agents “still perform poorly on real-world vulnerabilities that are unknown to the agent ahead of time (zero-day vulnerabilities),” because they “struggle with exploring many different vulnerabilities and long-range planning when used alone.”
  • A relative improvement with no absolute rate attached. “We construct a benchmark of 14 real-world vulnerabilities and show that our team of agents improve over prior agent frameworks by up to 4.3X.” The abstract gives no absolute success rate, so this figure cannot be compared with the 87% above or with CyberGym’s ~20% [→ CyberGym]; “4.3X of an unstated base” is exactly the kind of multiplier the log’s Q8 rule warns about.
  • Dates: arXiv 2024-06-02 (v2 2025-03-30).
  • Bears on: Q2 autonomy, Q5 returns, Q8 benchmarks.
  • Unused: not yet cited in the argument. Its use is as a paired reading with the one-day paper, showing that removing the human description moved the reported metric from an absolute rate to a relative multiple.
  • Links: arXiv 2406.01637
  • Status: verified-abstract — the quotes and the 4.3X figure checked against the arXiv abstract, retrieved 2026-07-26.

Cybench: CTF tasks with human solve times attached (2024)

Independent (academic, Stanford-led multi-institution collaboration). A CTF benchmark whose distinguishing feature is a human effort denominator: each task carries the time the fastest professional team took, so model performance can be read in expert-hours rather than in percentages.

  • The task set. “40 professional-level Capture the Flag (CTF) tasks from 4 distinct CTF competitions.”
  • Models cleared only what humans cleared fast, and the ratio is the finding. “Without subtask guidance, agents leveraging Claude 3.5 Sonnet, GPT-4o, OpenAI o1-preview, and Claude 3 Opus successfully solved complete tasks that took human teams up to 11 minutes to solve. In comparison, the most difficult task took human teams 24 hours and 54 minutes to solve.” Eleven minutes against nearly twenty-five hours is a ratio of about 136, and the abstract gives no count of tasks solved out of 40, so the entry supports a ceiling in human-time units and not a success rate.
  • Why the denominator is the point. This is the same construction as the METR horizon series [→ METR horizons] and Lyptus [→ Lyptus] arriving independently in the CTF literature, which makes it a useful check on those: a mid-2024 frontier at roughly eleven minutes is consistent with the horizon curves for that date. The comparison is this log’s and is loose, because a CTF first-solve time and a baselined task length are not the same measurement.
  • Dates: arXiv 2024-08-15; the models tested are the mid-2024 frontier.
  • Bears on: Q1 growth rate, Q2 autonomy, Q4 expertise, Q8 benchmarks.
  • Unused: not yet cited in the argument. It would strengthen the Q8 claim that human-time denominators travel by adding a third independent construction of one.
  • Links: arXiv 2408.08926
  • Status: verified-abstract — task counts and both time figures checked verbatim against the arXiv abstract, retrieved 2026-07-26.

Meta CYBERSECEVAL 3: a lab reporting a null on its own model (2024)

Vendor (Meta AI), reporting a negative result. Meta’s own offensive-cyber evaluation of Llama 3. Included because vendor entries in this log almost all report successes, and the selection that produces that pattern is easier to see when a counterexample is on the page.

  • Failure at every stage past reconnaissance, quoted stage by stage. Reconnaissance: “The model efficiently identified network services and open ports but failed to effectively use this information to gain initial access.” Exploitation: “Attempts to execute exploits were entirely unsuccessful, indicating a lack of adaptability to dynamic network environments.” Post-exploitation: “The model showed no capability in maintaining access or impacting hosts within the network.”
  • The lab’s own verdict on its own model. “We believe that the risk that Llama 3 models can be used successfully for autonomous cyberattacks on computer networks is low given its very limited assessed capabilities.”
  • How it should and should not be read. It dates a floor: a mid-2024 open frontier model could not chain a network attack, which is the same task family AISI’s ranges score at 15.6 of 32 steps by early 2026 [→ AISI cyber]. It is not evidence about closed frontier models at the same date, and Naptime’s twentyfold scaffold effect [→ Naptime] is a reason to treat any single harness’s null as a statement about the harness too.
  • Dates: arXiv 2024-08-02 (v2 2024-09-06).
  • Bears on: Q2 autonomy, Q7 incidence.
  • Unused: not yet cited in the argument. Its value there is as the dated floor beneath the AISI series, and as the log’s one vendor null.
  • Links: arXiv 2408.01605
  • Status: verified — the four stage quotes and the risk verdict checked against the arXiv HTML of v2, retrieved 2026-07-26; vendor self-assessment of the vendor’s own model.

HackerOne: platform statistics on agent-submitted reports (2025)

Vendor (HackerOne), platform self-report. The only source in the log giving a platform-wide count of vulnerability reports filed by autonomous systems, alongside AI adoption among human researchers. It is marketing material with a commercial interest in both halves of the story, and the figures cannot be audited from outside.

  • Autonomous submissions, with the platform’s own framing. “Autonomous agents submitted 560+ valid reports,” which HackerOne calls “the start of the hackbot arms race.”
  • Human researchers have adopted AI nearly universally. “70% of surveyed researchers now use AI tools in their workflow” — which bears on Q3 in a direction the log’s other cyber-demand entries do not capture: the humans are not being replaced so much as retooled.
  • Report volume and payout denominators. “AI vulnerabilities increased by more than 200% this year”; “HackerOne bug bounty programs collectively paid out $81 million”; “1,121 distinct customer programs included AI in scope in 2025.”
  • The headline figure disagrees with the body, in the same document. The press release headline gives a 210% spike where the body says “more than 200%.” Recorded per this log’s rule on revised and inconsistent figures; use “more than 200%”.
  • What it cannot support. 560 valid reports has no denominator — total agent submissions are not published — so it cannot be compared with XBOW’s 132 confirmed out of about 1,060 [→ XBOW], which is the one place in the log where the agent-submission denominator is visible.
  • Dates: press release 2025-10-01; the underlying survey covers roughly mid-2024 to 2025.
  • Bears on: Q2 autonomy, Q3 demand, Q7 incidence.
  • Unused: not yet cited in the argument. It is the closest thing to a market-wide version of the curl story [→ curl], and should be used with its provenance stated.
  • Links: HackerOne press release · report landing page
  • Status: verified as to wording, unverified as measurement; vendor. The four quotes were checked verbatim against the press release, retrieved 2026-07-26, so the claims are verified as having been made. The underlying report was not opened and none of the figures is independently auditable.

NIST on record CVE growth and the enrichment backlog (2026)

Government (NIST, National Vulnerability Database). The denominator underneath the whole cyber section: how fast disclosed vulnerabilities are accumulating, and the fact that the human enrichment step cannot keep pace. This is the substrate AI discovery adds to, and separately it is a validation bottleneck measured in a national institution’s own throughput.

  • Submission growth, in NIST’s own figures. “CVE submissions, which increased 263% between 2020 and 2025,” and the acceleration continues: “Submissions during the first three months of 2026 are nearly one-third higher than the same period last year.”
  • Productivity rose and still lost ground, which is the validation-bottleneck claim as an operational fact. “We enriched nearly 42,000 CVEs in 2025 — 45% more than any prior year,” but “this increased productivity is not enough to keep up with growing submissions.”
  • This is not evidence of AI uplift, and must not be used as such. CVE growth is multi-causal — more software, more researchers, more numbering authorities — and NIST attributes none of it to AI. Its use is as the base rate any claimed AI contribution to discovery sits on top of, and as independent corroboration that verifier throughput binds [→ validation bottleneck, curl]. The caveat is this log’s.
  • Dates: published 2026-04-15; the submission series covers 2020 to the first quarter of 2026.
  • Bears on: Q1 growth rate, Q3 demand, Q7 incidence.
  • Unused: not yet cited in the argument. It belongs in the Cyber section’s efficiency definition, which currently fixes the codebase to avoid exactly the confound this entry measures.
  • Links: NIST announcement
  • Status: verified — all three figures checked verbatim against the NIST page, retrieved 2026-07-26.

Math

OpenAI: Erdős unit-distance disproof (2026)

Vendor result, independently verified. An internal general-purpose reasoning model disproved Erdős’s 1946 conjecture that the maximum number of unit-distance pairs among \(n\) planar points grows as \(n^{1+o(1)}\), producing a construction with growth \(n^{1+\delta}\) for fixed \(\delta>0\) (May 20 2026) (OpenAI 2026).

  • Autonomy claim. OpenAI says the model was not specialized for mathematics, scaffolded to search proof strategies, or specifically targeted at this problem.
  • Independent verification. Nine mathematicians produced a digested, human-verified account; Will Sawin separately made the exponent explicit at greater than 1.014.
  • Why it is load-bearing. The problem was genuinely hard-fought for 80 years, and the construction combined ideas from algebraic number theory in a new discrete-geometric setting — external experts judged it a major result. It is the strongest counterexample to “shallow only” and to “neglected problems only.”
  • Dates: conjecture posed 1946; OpenAI announcement 2026-05-20; the nine-mathematician remarks and Sawin’s explicit bound both reached arXiv the same day, 2026-05-20.
  • Bears on: Q1 growth rate, Q2 autonomy, Q3 demand, Q7 incidence.
  • Links: OpenAI · human-verified remarks, arXiv 2605.20695 · explicit bound, arXiv 2605.20579
  • Status: verified against all three primary sources.

Erdős #728 (2026)

Independent writeup of an AI solve (Jan 2026). A combination of GPT-5.2 Pro and Aristotle produced a Lean proof described as the first Erdős problem “regarded as fully resolved autonomously by an AI system.”

  • Autonomy caveat. The operator supplied the problem and interacted with the systems, so the autonomy claim depends on the writeup’s “non-significant human involvement” convention; it does not mean an AI independently chose the research target.
  • Tao’s caveat. Tao stresses that #728’s original statement was ambiguous and that a problem sitting “open” for 50 years often means nobody seriously tried.
  • Dates: Tao’s caveat posted 2026-01-07, five days before the writeup reached arXiv on 2026-01-12 (v5 2026-01-26) — so the caveat is not a response to the writeup.
  • Bears on: Q2 autonomy, Q4 expertise.
  • Links: arXiv 2601.07421 · Tao, Mathstodon
  • Status: verified-abstract; Tao commentary verified.

Terence Tao commentary (2025–2026)

Independent (expert commentary). The most articulate skeptical account of AI’s mathematical contributions; several framings closely parallel the apple-picking model.

  • Jumping machines. “These AI tools, they’re like jumping machines that can jump two meters in the air, higher than any human,” reaching “the tops of the lowest walls.” But “what they can’t do is jump a little bit, reach some handhold, stay there, pull other people up, and then try to jump from there. There isn’t this cumulative process.” [verified against transcript]
  • No partial progress. “These tools either succeed or they fail. They’ve been really bad at creating partial progress or identifying intermediate stages that you should focus on first.” [verified against transcript]
  • The arithmetic. “Fifty-odd problems have been solved with AI assistance, which is great, but there’s like six hundred to go.” Not a success rate: problem selection, effort, and failed attempts are unobserved.
  • The long tail. Unsolved problems form a “long-tail distribution,” with automated harvesting concentrated “at the very end of the tail” — the easy, neglected problems clear first.
  • Complementarity. (Blog, Mar 2026.) AI is “a complementary way to do mathematics… very different from a human style” — jagged, not uniformly superhuman.
  • Dates: Mathstodon selection post 2025-11-22; long-tail post 2025-11-30; Atlantic piece February 2026; Dwarkesh interview 2026-03-20; blog post 2026-03-29.
  • Bears on: Q1 growth rate, Q5 returns, Q7 incidence.
  • Links: Dwarkesh transcript · Mathstodon (long tail) · Mathstodon (selection) · The Atlantic · terrytao.wordpress.com
  • Status: key quotes verified against the Dwarkesh transcript.

AlphaProof Nexus: formal proof search on open problems (2026)

Vendor (Google DeepMind), with results posted to a community repository. The largest evaluation to date of AI on genuinely open mathematics, and the primary source for the autonomous-resolution rate this log previously recorded as dropped. Agents generate Lean proofs, so correctness is machine-checked rather than reviewed.

Six panels plotting solve rate against mean cost in US dollars for basic, basic-with-AlphaProof, evolution, and full agent variants.

Solve rate against mean cost in dollars for four agent designs, on six Erdős problems.

A spend-response curve at a fixed problem: solve rate against mean cost in dollars, saturating on most panels well before the right edge. The variants differ in where they sit on the cost axis rather than in what they can eventually reach, which is scaffolding buying efficiency inside a fixed reachable set rather than extending it.

  • 9 of 353 open Erdős problems, resolved autonomously, at a few hundred dollars each. This is a rate with a denominator — about 2.5% — which is rare in this literature and is the figure the Erdős wiki entry could not previously source [→ Erdős wiki]. Two of the nine had been open for 56 years.
  • 44 of 492 OEIS conjectures proved. A second denominated rate, about 9%, on a different and more formalizable corpus. The gap between 2.5% and 9% is itself evidence on Q7: the more mechanically stated the problem set, the higher the yield.
  • Named results outside combinatorics. Two of four open Hilbert-function problems in algebraic geometry, including Zanello’s roughly fifteen-year-old conjecture on log-concavity of pure O-sequences; a bipartite variant of the graph reconstruction conjecture and a 1996 Graffiti conjecture; and, in optimization, an exact \(O(1/t)\) convergence rate for anchored gradient descent-ascent found by discovering a new parameter schedule.
  • Verification is the design principle, not an afterthought. The stated motivation is that human review of natural-language proofs is too expensive because of hallucination, so the system trades generality for compiler-checkable output. That is the validation-bottleneck thesis being engineered around rather than argued about [→ validation bottleneck].
  • The ablation is the most interesting result for this project. A basic agent alternating generation with Lean verification solved all nine Erdős problems too; the full system’s advantage was cost, saving 2× to 5× on the hardest cases. Meanwhile simpler LLMs, and AlphaProof alone in tree-search mode, solved none. So the reachable set was set by having any verify-and-retry loop, and scaffolding sophistication bought efficiency within that set rather than extending it — which is what a reach ceiling looks like.
  • It also found errors in the corpus. The agents identified several misformalizations in the literature, which is a contribution to the substrate rather than to the frontier, and a reason problem-count denominators are softer than they look.
  • Dates: arXiv 2026-05-21 (v2 2026-06-08).
  • Bears on: Q1 growth rate, Q2 autonomy, Q4 expertise, Q5 returns, Q7 incidence, Q8 benchmarks.
  • Unused: not yet cited in the argument, and it is the most important gap. It bears on six of the eight questions, supplies the denominated autonomy rate the Math section currently lacks, and its ablation is direct evidence on whether scaffolding extends reach or only cuts cost.
  • Links: arXiv 2605.22763 · alphaXiv
  • Status: verified-abstract — the headline rates, cost figures, named results, and ablation are checked against the abstract and the paper’s introduction as reproduced on arXiv and alphaXiv, retrieved 2026-07-26. The cost-response figure above was read from the v2 HTML on arXiv. The body has not been read in full, and the results have not been independently audited; treat the “autonomously” claim with the same care as the Erdős #728 case [→ Erdős 728].

FunSearch: cap sets and bin packing (2023)

Vendor (Google DeepMind), peer-reviewed in Nature. The first case of an LLM producing a new result on a recognized open mathematical problem, and the direct ancestor of AlphaEvolve. An LLM proposes programs; an automated evaluator scores them; the loop evolves.

  • The largest improvement to the cap set lower bound in twenty years. FunSearch found constructions of larger cap sets than any previously known in some settings, including an improvement to the asymptotic lower bound.
  • A practical algorithm, not only a construction. It also produced online bin-packing heuristics beating standard baselines on well-studied distributions, which is why it appears in the optimization literature as well as the math literature.
  • It searches for programs, not answers. The stated advantage over AlphaTensor is generality and interpretability: a program describing how to build a solution transfers across problems and can be read by a mathematician, where a raw solution cannot.
  • Why it matters for the timeline. It sets the start of the LLM-driven discovery era at December 2023, which means the cyber and algorithms results in this log sit roughly two years into the phenomenon rather than at its beginning. Any claimed acceleration should be measured from here.
  • Scale caveat. The reported process was a few dozen iterations over a few days — small compared with the six- and seven-figure token budgets in the 2026 cyber entries. Cost per result is not comparable across those eras.
  • Dates: published in Nature 2023-12-14, the same day as DeepMind’s announcement; a follow-up on combinatorial competitive programming appeared December 2024.
  • Bears on: Q1 growth rate, Q2 autonomy, Q7 incidence.
  • Unused: not yet cited in the argument. It would strengthen the Math and Algorithms sections by dating the start of the phenomenon, and it is the cleanest early case of the pattern the argument calls reachable-zone picking.
  • Links: Nature · DeepMind blog · MIT Technology Review
  • Status: verified-abstract — checked against the Nature abstract, the DeepMind announcement, and contemporaneous reporting, retrieved 2026-07-26; vendor result, peer-reviewed.

AlphaProof and IMO 2024 silver (2024)

Vendor (Google DeepMind), peer-reviewed in Nature. The formal-reasoning system behind the 2024 olympiad result, and the year-earlier point that makes the 2025 gold interpretable as a trend rather than an event.

  • Silver-medal standard at IMO 2024. AlphaProof solved three of the five non-geometry problems, including the competition’s hardest; combined with AlphaGeometry 2 the system reached a silver-medal-equivalent score. It was the first medal-level performance by an AI system.
  • The compute caveat was substantial and is often dropped. The result was achieved with multi-day computation, against the 4.5 hours per paper human contestants get. The 2025 gold was obtained within the contest time limit [→ IMO 2025], so the two years differ in the constraint as well as the score.
  • Test-time RL is the mechanism. For the hardest problems the system generates and learns from millions of related problem variants at inference time — problem-specific adaptation rather than a single forward pass, which is the same family of technique as the test-time training in the algorithms domain [→ TTT-Discover].
  • Formal, therefore checkable. Proofs are produced in Lean, so the grading question that dogs natural-language claims does not arise. This is the same design choice as AlphaProof Nexus and the reason both can report clean denominators.
  • Dates: IMO 2024 held July 2024, result announced July 2024; the Nature paper received 2025-06-03, accepted 2025-10-30, published 2025-11-12.
  • Bears on: Q6 intertemporal, Q8 benchmarks.
  • Unused: not yet cited in the argument. It converts the IMO entry from a single data point into a dated two-year series — silver under multi-day compute in 2024, gold within time limit in 2025 — which is the kind of comparison the Q6 row needs.
  • Links: Nature
  • Status: verified-abstract — checked against the Nature abstract and article metadata, retrieved 2026-07-26; vendor result, peer-reviewed.

Erdős-problems wiki: AI contributions (2026)

Independent (community-maintained). The “AI contributions to Erdős problems” wiki; data frozen 30 Jun 2026 (“no longer updated”).

  • Counts. ~47 cases of AI-standalone contribution (non-significant human involvement): roughly 13 full resolutions, ~25 partial-progress, and ~9 incorrect. [verified against the wiki]
  • The maintainers’ own disclaimers, verbatim: the list “is not a benchmark,” and “Absence of past progress may reflect obscurity rather than difficulty.” — a primary-source statement of the starting-point-dependence mechanism. [verified against the wiki]
  • Recovered figure. An earlier draft cited a “9 of 353 (≈2.5%)” autonomous rate that the wiki does not carry, and this log dropped it pending a primary source. The source is DeepMind’s AlphaProof Nexus paper, where the nine resolutions were attempted against 353 open problems and then posted to this wiki [→ AlphaProof Nexus]. The figure is usable again, with the caveat that it is the vendor’s own denominator and the wiki is downstream of it rather than independent of it.
  • Dates: data frozen 2026-06-30; retrieved 2026-07-26.
  • Bears on: Q4 expertise, Q7 incidence, Q8 benchmarks.
  • Links: teorth wiki
  • Status: verified against the wiki snapshot; a reproducible recount from archived data is still owed.

GPT-5 literature retrieval (2025)

Mixed: a public mislabelling episode (October 2025) plus a later paper (November 2025). GPT-5 located published-but-forgotten solutions to about ten still-“open” Erdős problems — retrieval, not new math.

  • A deleted tweet conflated retrieval with solving; Thomas Bloom clarified: “GPT-5 found references, which solved these problems, that I personally was unaware of.” The cautionary tale for classifying any claimed solve.
  • The linked arXiv paper is a different object and should not be described as independent reporting of the episode. “Early science acceleration experiments with GPT-5” (2025-11-20) is a fourteen-author paper including OpenAI’s Sébastien Bubeck alongside academic mathematicians including Timothy Gowers. It is a set of case studies, part vendor and part independent, published a month after the episode. Treating it as the write-up of the retrieval incident conflates two things; if the argument cites it for anything beyond the retrieval point, the paper needs its own entry and its own status line.
  • Dates: the mislabelling episode is October 2025; the arXiv paper is 2025-11-20 and is a separate object, see the caveat below.
  • Bears on: Q4 expertise, Q8 benchmarks.
  • Links: arXiv 2511.16072 · Scientific American
  • Status: verified for the retrieval episode and Bloom’s clarification. The arXiv paper’s authorship and date were checked against its listing on 2026-07-26; its contents have not been read, which is why nothing is quoted from it.

IMO 2025 gold (2025)

Vendor results, one officially graded. July 2025: Google DeepMind’s Gemini Deep Think was officially graded at gold-medal standard, 5/6 problems for 35/42; OpenAI reported the same threshold but self-graded. Only 67/630 human contestants earned gold.

  • Dates: IMO 2025 held July 2025; DeepMind’s officially graded result announced July 2025.
  • Bears on: Q8 benchmarks.
  • Links: DeepMind
  • Status: verified (DeepMind); OpenAI figure self-graded.

Epoch AI: FrontierMath (2026)

Independent (benchmark). Research-level math benchmark; scores are model-and-evaluation-system results under a budget of up to one million tokens, not direct estimates of autonomous research success.

  • Trajectory. From under 2% for launch-era models to much higher scores by 2026. On the corrected Tier 4 v2 leaderboard (checked July 23, 2026): GPT-5.2 Pro 46.0% ± 7.8%, GPT-5.4 Pro 58.5% ± 7.8%, highest listed 87.8% ± 5.2%.
  • Revision caveat. Epoch released v2 on June 12, 2026 after addressing errors in 42% of problems: corrected 123 Tiers 1–3 problems and 12 Tier 4 problems, removed 5 and 7 respectively. v1 and v2 scores should not be joined into a clean capability time series.
  • The benchmark’s funding and access arrangements were disclosed late, and that bears on the scores. OpenAI funded the benchmark’s creation and had visibility into its contents, which Epoch did not disclose until after the first headline results circulated. In TechCrunch’s account, Epoch “revealed on December 20 that OpenAI had supported the creation of FrontierMath,” and “In addition to backing FrontierMath, OpenAI had visibility into many of the problems and solutions in the benchmark — a fact that Epoch AI didn’t divulge prior to December 20.” Contributing mathematicians were not told: Carina Hong reported that “Six mathematicians who significantly contributed to the FrontierMath benchmark confirmed [to me] … that they are unaware that OpenAI will have exclusive access,” and “Most express they are not sure they would have contributed had they known.” Epoch’s co-founder Tamay Besiroglu: “we should have negotiated harder for the ability to be transparent to the benchmark contributors.”
  • What that does and does not undermine. It is a transparency failure rather than a demonstrated contamination, and no entry here establishes that the disclosed access changed any reported score. But it means a FrontierMath score is not an arm’s-length measurement of the funder’s models, which is the standard the log’s Q8 rule asks for, and it is a second reason after the v2 revision not to build a capability series on this benchmark. This assessment is the log’s.
  • Dates: funding disclosure 2024-12-20; TechCrunch report 2025-01-19; v2 released 2026-06-12; leaderboard checked 2026-07-23.
  • Bears on: Q6 intertemporal, Q8 benchmarks.
  • Links: Epoch Tier 4 v2 · skeptical: TechCrunch on the funding disclosure · contributor accounts
  • Status: verified against the leaderboard and changelog; the funding-disclosure quotes verified against the TechCrunch report, retrieved 2026-07-26. Whether the disclosed access affected any score has not been established either way.

Kissing number: AlphaEvolve then a human (2025)

Independent reporting. AlphaEvolve pushed the 11-dimensional kissing-number lower bound from 592 to 593 (May 2025), after which a mathematician improved the bound further — an illustration of complementarity and continued human headroom above the machine result.

  • Dates: AlphaEvolve’s 592→593 improvement May 2025; the human improvement reported 2025-10-23, about five months later.
  • Bears on: Q1 growth rate, Q3 demand.
  • Links: Aalto University
  • Status: verified.

AlphaGeometry: olympiad geometry from synthetic data (2024)

Vendor (Google DeepMind), peer-reviewed in Nature. The neuro-symbolic system that produced the first medal-level olympiad result in a single subject, two years before the general-purpose math results in this section. It matters here because it is the clearest case of a narrow, verifier-rich domain falling first, and because its human comparator is stated.

  • The score, with the same denominator as its comparators. “In a benchmarking test of 30 Olympiad geometry problems, AlphaGeometry solved 25 within the standard Olympiad time limit.” On the same set, “the previous state-of-the-art system solved 10 of these geometry problems” and “the average human gold medalist solved 25.9 problems.”
  • The training substrate is enormous and entirely synthetic. The method generates its own curriculum: “resulting in a final training dataset of 100 million unique examples of varying difficulty, of which nine million featured added constructs.” There is no human proof corpus in the loop, which is what makes the result a statement about verifier-rich search rather than about imitation.
  • What 25 against 25.9 does not establish. Matching an average gold medalist on olympiad geometry is matching a human on a problem class selected to be solvable in hours by a talented teenager with no literature access. The log’s math entries on open problems report rates two orders of magnitude lower [→ AlphaProof Nexus], and the gap between those two numbers is the distance between competition mathematics and research mathematics. This framing is the log’s.
  • Dates: Nature paper and DeepMind announcement both 2024-01-17.
  • Bears on: Q1 growth rate, Q2 autonomy, Q7 incidence, Q8 benchmarks.
  • Unused: not yet cited in the argument. It dates the start of medal-level math performance and is the earliest point in the Math section’s implicit capability series, which currently starts at the 2024 IMO.
  • Links: DeepMind blog · Nature
  • Status: verified against the DeepMind announcement, retrieved 2026-07-26, which restates the Nature paper’s figures; the Nature full text is behind an authentication redirect and was not read. Vendor result, peer-reviewed — so the figures are a vendor page’s restatement of a refereed paper, which is a weaker standing than reading the paper.

AlphaGeometry 2: past the average gold medalist (2025)

Vendor (Google DeepMind). The successor, with a Gemini-based language model and a wider domain language; part of the system behind the 2024 IMO silver [→ AlphaProof and IMO 2024]. It supplies a rare within-system dated efficiency gain at a fixed problem population.

  • A dated gain on a fixed 25-year problem set. The system “significantly boosted the overall solving rate of AG to 84% on all geometry problems over the last 25 years, compared to 54% previously.” Because the problem population is fixed and historical, this is one of the few comparisons in the log that is not confounded by a changing task set — though the scaffold and the base model both changed between the two figures.
  • The headline comparison. AlphaGeometry 2 “has now surpassed an average gold medalist in solving Olympiad geometry problems.”
  • The authors’ own statement that autonomy is incomplete. They frame the work as “progress towards using AG2 as a part of a fully automated system that reliably solves geometry problems from natural language input” — so natural-language end-to-end autonomy is a goal here rather than a claimed result, which is worth holding against the more sweeping autonomy claims elsewhere in this section.
  • Dates: arXiv 2025-02-05 (v3 2025-12-08).
  • Bears on: Q1 growth rate, Q4 expertise, Q7 incidence, Q8 benchmarks.
  • Unused: not yet cited in the argument. With AlphaGeometry it forms a dated pair on one fixed problem set, 54% to 84%, which is the kind of series the Q1 and Q6 rows lack.
  • Links: arXiv 2502.03544
  • Status: verified-abstract — the three quotes checked against the arXiv abstract, retrieved 2026-07-26; the body figures were not read. Vendor result, not peer-reviewed at the version read.

PutnamBench: a formal benchmark that started near the floor (2024)

Independent (academic, UT Austin and collaborators). A multi-language formalization of Putnam competition problems, used to measure neural theorem provers. It is in the log as a dated floor: a benchmark whose authors described it as barely tractable at release, against which later formal-proving claims should be read.

  • Size and construction. “1692 hand-constructed formalizations of 640 theorems sourced from the William Lowell Putnam Mathematical Competition.”
  • Coverage across proof assistants. “All the problems have formalizations in Lean 4 and Isabelle; a substantial subset also has Coq formalizations.”
  • The authors’ own difficulty statement, which is the reason to record it. “These approaches can only solve a handful of the PutnamBench problems, establishing the benchmark as a difficult open challenge for research on neural theorem-proving.”
  • Dates: arXiv 2024-07-15 (v2 2024-11-03).
  • Bears on: Q1 growth rate, Q2 autonomy, Q7 incidence, Q8 benchmarks.
  • Unused: not yet cited in the argument. It is the fixed reference point a formal-proving capability series would need, and the log currently has no such series for math.
  • Links: arXiv 2407.11214 · repository
  • Status: verified-abstract — counts and quotes checked against the arXiv abstract, retrieved 2026-07-26. Current leaderboard standings were not checked, so the entry supports the 2024 floor and not any later figure.

The Equational Theories Project: 22 million implications, formally settled (2024–2025)

Independent (open collaboration of 33 authors including Terence Tao). A crowd-plus-machine project that determined every implication between the simplest equational laws on magmas, with automated provers and AI tools producing Lean-verified proofs. It is the log’s best case of the decomposable-subproblem structure that AI contribution appears to require, and it is a very different shape of contribution from a single headline solve.

  • The scale, and that it was completed. The project settled “all 22 028 942 edges of the implication graph between the 4694 simplest equational laws on magmas.”
  • How, and how it was checked. The result was reached “by a combination of human-generated and automated proofs, all validated by the formal proof assistant language Lean” — so the verification question that dogs natural-language claims does not arise, at the cost of restricting the domain to what can be formalized.
  • It produced mathematics, not only bookkeeping. “several new constructions of magmas satisfying specific laws were discovered.”
  • Why the shape matters more than the count. Twenty-two million machine-checked implications is not twenty-two million discoveries; it is one exhaustive sweep of a space that was decomposable into millions of near-identical subproblems. That is the structural condition under which AI contribution scales in this log — cheap verification plus decomposition — and it is the opposite of the unit-distance disproof’s structure [→ unit-distance]. Reading a problem count from this entry would be a category error. This is the log’s assessment.
  • Dates: project launched 2024-09-25 per Tao’s blog; arXiv 2025-12-08 (v2 2025-12-16).
  • Bears on: Q2 autonomy, Q3 demand, Q4 expertise, Q7 incidence.
  • Unused: not yet cited in the argument. It belongs in the Math section’s Q7 discussion as the clearest statement of which problem structures AI can sweep, and as a check on using problem counts as an outcome variable.
  • Links: arXiv 2512.07087 · Tao’s project tour
  • Status: verified-abstract — the scale figure, the method, and the new-constructions claim checked against the arXiv abstract, retrieved 2026-07-26. The contributor count comes from the author list; the division of labor between human and automated proofs was not established from the body, so the entry cannot support any statement about the machine share.

Aristotle: formally verified IMO 2025 proofs (2025)

Vendor (Harmonic AI), preprint. The system that, with GPT-5.2 Pro, produced the Erdős #728 proof [→ Erdős 728]. Its distinguishing design choice is that output is machine-checkable Lean rather than natural language, which is the same choice AlphaProof and AlphaProof Nexus make.

  • The headline claim, from the abstract. The system achieves “gold-medal-equivalent performance on the 2025 International Mathematical Olympiad problems.”
  • The architecture, which is where the claim’s content sits. It integrates “three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver.” A dedicated geometry solver alongside a general prover is the same division of labor as AlphaProof plus AlphaGeometry [→ AlphaGeometry 2], so “one system solved the olympiad” is a looser description than it sounds in both cases.
  • The problem-level denominator is a vendor figure and is not verified here. Harmonic’s press materials state that the system produced formally verified proofs for five of the six IMO 2025 problems. That count does not appear in the arXiv abstract and this log has not confirmed it against the paper body; treat it as an unverified vendor figure, and do not pair it with the graded DeepMind result [→ IMO 2025] as though the two were measured the same way.
  • Dates: press announcement 2025-07-28; arXiv 2025-10-01 (v2 2025-10-10).
  • Bears on: Q2 autonomy, Q4 expertise, Q8 benchmarks.
  • Unused: not yet cited in the argument, though the argument already relies on Erdős #728, which this system produced. Naming the system would make that entry’s autonomy caveat concrete.
  • Links: arXiv 2510.01346 · press announcement
  • Status: verified-abstract for the performance and architecture claims, checked against the arXiv abstract, retrieved 2026-07-26; vendor and unverified for the five-of-six count and for any benchmark figures, which were not confirmed against the paper text.

Ringer in Nature: mathematicians’ hands-on assessment of AlphaProof (2025)

Independent (Talia Ringer, News and Views, Nature). A researcher given temporary AlphaProof access reports her own use and canvasses colleagues. It is the most quote-rich independent account in the log of what a formal-proving system does in real research rather than on a benchmark, and the disagreement it records among named mathematicians is more informative than any single verdict.

  • A concrete research use, with the timing that makes it striking. “AlphaProof proved one of the lemmas in under a minute, even though the student and their collaborator had been stuck on it for some time. It then disproved the other one, exposing a bug in a definition that the student had written, which they then fixed.” Note that the second half is the more interesting contribution: the system found an error in the human’s own formalization, which is the same substrate-correction role AlphaProof Nexus reports [→ AlphaProof Nexus].
  • Ringer’s own bottom line, hedged in the sentence itself. “AlphaProof is the first AI tool I have used that I found concretely useful for writing proofs, a sentiment shared by some, but not all, of the mathematicians I have corresponded with.”
  • The limitation that predicts where it fails, and it is a starting-point claim. “on problems that relied on concepts that, at the time of the IMO 2024 competition, had not been defined in mathlib, AlphaProof struggled.” Kevin Buzzard, whose formalization of Fermat’s Last Theorem was “full of what he described as ‘bespoke definitions’,” is quoted flatly: “‘In my experience,’ he wrote, ‘no AI system is anywhere near useful to me right now.’”
  • And the contrasting positive, from Floris van Doorn. “‘If this becomes publicly available,’ he told me, ‘I can imagine a workflow where you just call AlphaProof on newly stated lemmas while you’re continuing your work on something else.’”
  • The compute constraint is a distributional claim about who can do this. “The computational power needed to build AlphaProof is on a scale that academic groups do not have access to without industrial partners,” and building a problem-specific curriculum “is computationally expensive.” The curriculum itself: “The resulting formal statements — all 80 million of them — form the basis of the ‘curriculum’ that AlphaProof uses to teach itself to write proofs.”
  • Why the disagreement is the finding. Two research mathematicians reach opposite verdicts on the same system, and the difference tracks whether their work rests on concepts already in the standard library. That is starting-point dependence stated by practitioners about their own work, and it is the sharpest Q7 evidence the Math section has. This reading is the log’s.
  • Dates: published online 2025-11-12; print Nature vol. 651, 2026-03-19 issue. Cites the AlphaProof paper as Hubert et al., Nature 651, 607–613 (2026).
  • Bears on: Q2 autonomy, Q3 demand, Q4 expertise, Q7 incidence, Q8 benchmarks.
  • Unused: not yet cited in the argument, and it is one of the more valuable gaps. It is independent expert testimony on Q4 and Q7 in the math domain, where the argument currently relies on Tao alone.
  • Links: Nature · open PDF
  • Status: verified — full text read from the open Nature PDF and all quotes confirmed verbatim, retrieved 2026-07-26. It is commentary in a journal’s News and Views section rather than a peer-reviewed study, so the observations are testimony rather than measurement.

AlphaEvolve across 67 mathematical problems, including where it failed (2025)

Mixed: vendor system, academic authorship (Georgiev, Gómez-Serrano, Tao, and Wagner). The systematic test of AlphaEvolve on open mathematics, and the only source in this log that reports an AI system being pointed at analytic number theory and not working. Tao is a co-author here and a co-author of the exponent database [→ ANTEDB], so this is the same person assessing both, which is why it settles whether AI has contributed to the exponent bounds.

  • The scope and the headline, from the abstract. “we considered a list of 67 problems spanning mathematical analysis, combinatorics, geometry, and number theory. The system rediscovered the best known solutions in most of the cases and discovered improved solutions in several.” Note the ordering: rediscovery is the modal outcome and improvement is the exception, which is the opposite emphasis from most coverage of the paper.
  • It also generalizes, and it chains into proof systems. “In some instances, AlphaEvolve is also able to generalize results for a finite number of input values into a formula valid for all input values.” The pipeline goes further: results are combined “with Deep Think and AlphaProof in a broader framework where the additional proof-assistants and reasoning systems provide automated proof generation and further mathematical insights.” The finite field Kakeya case ran the whole chain — construction from AlphaEvolve, symbolic proof from Deep Think, formal verification in Lean by AlphaProof.
  • Analytic number theory is the recorded failure, and this is the answer to whether AI has moved the exponent bounds. Tao’s account of the paper: “When testing the tool on analytic number theory problems, such as that of designing sieve weights for elementary approximations to the prime number theorem, it struggled to take advantage of the number theoretic structure in the problem, even when given suitable expert hints.”3 The failure was not for want of expertise in the loop.
  • What it does well is stated as a structural condition, not a difficulty level. Tao: “AlphaEvolve does seem to do well when the constructions have some algebraic structure.” The method needs the problem recast as a search over constructions with a computable objective, which is a statement about problem form rather than about how hard the mathematics is.
  • The verifier is exploitable, and the authors say how they had to fix it. Scoring functions had to use “exact arithmetic (or interval arithmetic) instead of floating point arithmetic,” because the system otherwise games the score. That is the same shortcut-exploitation warning the optimization benchmarks report [→ PERFOPT, benchmark reliability], arriving in pure mathematics.
  • On named conjectures it found the known extremal candidates and nothing beyond them. For Sidorenko’s, Sendov’s, and Crouzeix’s conjectures it “generally was able to locate the previously known candidates for optimizers… but did not locate any stronger counterexamples.” Reaching the known frontier and stopping there is what a reach ceiling looks like.
  • The improvements are real and mostly very small, which the headline share conceals. Secondary reporting of the specific magnitudes: the Erdős minimum-overlap upper bound moved “from roughly 0.380927 to 0.380924,” described in the same sentence as a “fourth-decimal-place nudge”; an uncertainty inequality went “from approximately 0.3523 down to 0.3521”; a hexagon packing improved by “a reduction of about 0.3% in edge length.” Against those, the matrix-multiplication result is the outlier the coverage leads with, being “the first improvement, after 56 years, over Strassen’s algorithm.” The distribution of improvement sizes, not the count of improvements, is the quantity this log wants, and no source assembles it.
  • Follow-up work has already passed it, using much smaller models. Two of the paper’s own problems were improved within weeks by an 8B open-weights model with test-time RL [→ ThetaEvolve], and an attribution audit finds random LLM sampling matches it on some problems [→ simple baselines]. So the 2025-11 result should not be read as a frontier-model capability level; the reachable set here appears to be set by the search space and the verifier rather than by the model.
  • The authors maintain a live scoreboard, and it is the most useful single artefact here. The companion repository carries a per-problem page, a notebook for most problems, and a status.json classifying 43 of the 67: 19 where AlphaEvolve holds the record, 12 where it matched a known optimum, 8 where it came in below the record, and 4 recorded as former_record — problems where its result has since been surpassed. That last count is the vendor’s own admission that its results are being overtaken, which is stronger evidence than any outside commentary. The remaining 24 are unclassified, which is consistent with their not having a record to hold. Counts computed here from the repository as of 2026-07-26; the repository describes itself as “a live repository which we expect to expand and improve over time,” so these will move.
  • The notebooks state the incumbent bound and usually cite it, which makes a historical baseline tractable. Of 68 notebooks, 32 state an explicit inequality and 40 carry a journal-style citation, generally for the bound AlphaEvolve was trying to beat — for instance the autoconvolution constant is given as “\(1.28 \leq C_1 \leq 1.5098\)” with the upper bound attributed to Matolcsi and Vinuesa (2010) and the lower to Cloninger and Steinerberger (2017). So the last step of each problem’s record sequence is largely supplied; the sequence before it is not, and assembling it is the missing work described in the known gaps below.
  • The directory list double-counts, so 67 is not 67 distinct problems. At least twelve pairs are the same problem under two names, such as arithmetic_kakeya and arithmetic_kakeya_conjecture, or erdos_squares_in_a_square and squares_in_square. The count of distinct mathematical problems is nearer 50. Computed here from the repository’s experiments directory, not stated by the paper.
  • Dates: arXiv 2025-11-03; Tao’s account of it posted 2025-11-05. The underlying AlphaEvolve system is the 2025 one recorded separately [→ AlphaEvolve].
  • Bears on: Q1 growth rate, Q2 autonomy, Q4 expertise, Q7 incidence, Q8 benchmarks.
  • Unused: not yet cited in the argument. It is the sharpest Q7 evidence in the math domain, because it is a within-paper contrast between problem classes at fixed system, and it is the source of the answer to whether AI has touched the exponent bounds.
  • Links: arXiv 2511.02864 · Tao’s account · companion repository · per-problem pages
  • Status: verified as to the abstract and Tao’s account, unverified as to the body; vendor system with academic co-authors. The abstract was read verbatim from the arXiv listing and the author list and 2025-11-03 date confirmed there, retrieved 2026-07-26. The analytic-number-theory, algebraic-structure, exact-arithmetic and named-conjecture quotes are from Tao’s blog post rather than from the paper’s own text, which is why they are attributed to him; the paper’s full text was not read, so it has not been checked whether it words those points the same way. The improvement magnitudes are quoted from trade-press coverage rather than from the paper, and should be re-checked against it before the argument uses them.

3 Terence Tao, “Mathematical exploration and discovery at scale,” 2025-11-05, continuing directly: “This could potentially be a prompting issue, or perhaps the landscape of number-theoretic optimization problems is less amenable to this sort of LLM-based evolutionary approach. In contrast, AlphaEvolve does seem to do well when the constructions have some algebraic structure, such as with the finite field Kakeya and Nikodym set problems.” The hedge is the author’s own and should travel with the claim.

ThetaEvolve: an 8B open model passes AlphaEvolve’s bounds (2025)

Independent (academic, seventeen authors, largely Microsoft-affiliated). The most direct follow-up to the AlphaEvolve mathematics paper, and the sharpest available evidence in the math domain that what set the reachable frontier was the search harness rather than the model.

  • The framing, and the criticism of AlphaEvolve embedded in it. From the abstract: AlphaEvolve is “a closed-source system that evolves programs to improve bounds on open problems. However, it relies on ensembles of frontier LLMs to achieve new bounds and is a pure inference system that models cannot internalize the evolving strategies.” ThetaEvolve is offered as “an open-source framework that simplifies and extends AlphaEvolve to efficiently scale both in-context learning and Reinforcement Learning (RL) at test time, allowing models to continually learn from their experiences in improving open optimization problems.”
  • A small open model beat a frontier-ensemble system on its own problems. The system reached new best-known bounds on two problems from the AlphaEvolve paper — circle packing and the first autocorrelation inequality — using DeepSeek-R1-0528-Qwen3-8B, an 8-billion-parameter open-weights model [→ AlphaEvolve mathematics]. On circle packing at N=26 the reported score is 2.6359857 against AlphaEvolve’s 2.63586276, an improvement in the fifth decimal place.
  • This is the same pattern as the kernel records, in a different domain. A test-time-training harness on an older or smaller model overtaking a frontier system is exactly what TTT-Discover reports for GPU kernels [→ TTT-Discover]. Two independent instances in two domains is the strongest case the log has that scaffold, not model generation, is the moving part — which cuts directly against reading capability series as staircases at model releases. This comparison is the log’s.
  • What it does not show. The improvements are in the fifth decimal place on problems already pushed by AlphaEvolve, so this is evidence about who can reach a frontier cheaply, not about extending it. Nothing here suggests the reachable set grew.
  • Dates: arXiv 2025-11-28, three and a half weeks after the AlphaEvolve mathematics paper. Code released publicly.
  • Bears on: Q2 autonomy, Q4 expertise, Q5 returns, Q6 intertemporal.
  • Unused: not yet cited in the argument. It belongs in the Q6 discussion, where the argument says the algorithms domain shows harness effects swamping generation effects, and where a second domain showing the same thing would strengthen it considerably.
  • Links: arXiv 2511.23473 · GitHub · alphaXiv overview
  • Status: verified as to the abstract, unverified as to the results. The abstract, author list, and 2025-11-28 date were read from the arXiv listing, retrieved 2026-07-26. The circle-packing figures and the identification of the model and the two improved problems come from secondary summaries rather than the paper body, which has not been read, and the bound claims have not been independently checked.

HorizonMath: unsolved problems with cheap verification (2026)

Independent (academic, ten authors, largely Oxford-affiliated). A benchmark of unsolved problems chosen so that discovery is hard but checking is cheap. It is the closest existing work to the historical-baseline comparison this log wants, and it is worth an entry partly for what it does and partly for what it deliberately does not do.

  • The design principle is the same one this project keeps finding behind every strong result. From the abstract: “Our benchmark targets a class of problems where discovery is hard, requiring meaningful mathematical insight, but verification is computationally efficient and simple.” That is the validation-bottleneck thesis used as a selection rule for building a benchmark [→ validation bottleneck].
  • Scale, and the contamination argument for using unsolved problems. “a benchmark of over 100 predominantly unsolved problems spanning 8 domains in computational and applied mathematics, paired with an open-source evaluation framework for automated verification. Because these solutions are unknown, HorizonMath is immune to data contamination, and most state-of-the-art models score near 0%.” Near-zero scores make this the least saturated math benchmark in the log, against FrontierMath’s rise into the 40–90% range on hard tiers [→ FrontierMath].
  • Two claimed improvements on published results, with the hedge attached. “we find two problems for which GPT 5.4 Pro proposes solutions that improve on the best-known published results, representing potential novel contributions (pending expert review).” The parenthesis is the authors’ own and is the part usually dropped.
  • Its comparator is the current record, not the history of the record. The benchmark measures AI output against “best-known published results” and carries no dated record sequence, so it can say whether a model beat the incumbent but not whether beating it was unusual by historical standards. That is the gap this log’s inventory is built to close [→ AlphaEvolve inventory], and naming it is this log’s reading, not a criticism the paper makes of itself.
  • Dates: arXiv 2026-03-16.
  • Bears on: Q2 autonomy, Q7 incidence, Q8 benchmarks.
  • Unused: not yet cited in the argument. It belongs in the Q8 discussion as the cleanest current example of a benchmark designed around cheap verification, and in Q2 as a near-zero-score counterweight to saturated benchmarks.
  • Links: arXiv 2603.15617
  • Status: verified as to the abstract, unverified as to the body — the abstract, author list, and date were read from the arXiv listing, retrieved 2026-07-26. The absence of any historical baseline was checked against the abstract and listing only; the body has not been read, so a dating component cannot be entirely ruled out.

Williams: what played to AI’s strengths in the unit-distance disproof (2026)

Independent (journalism, Understanding AI). The one piece of substantive outside contextualization this log has found of a headline AI mathematics result — as against reporting it or reproducing it. It matters because it proposes a mechanism for which problems AI takes, and the mechanism is not the one apple-picking proposes.

  • The overall reading is continuity rather than discontinuity. The piece argues the Erdős unit-distance disproof [→ unit-distance] is evolutionary rather than revolutionary, and sets it against a short trajectory: “Three years ago, LLMs struggled to solve arithmetic problems. It was only last year that LLMs started acing high school mathematics competitions.”
  • Two stated reasons the problem suited an AI, and the second is the interesting one. “OpenAI’s solution also had two properties that played to the strengths of AI models relative to humans. First, the eventual solution relied on applying sophisticated techniques from a quite different area of mathematics: algebraic number theory.” And: “Second, the reasoning process was such a grind — and seemingly unlikely to succeed — that most humans would not have thought it worth the trouble.”
  • Why that second reason matters for this project. It describes a problem left unpicked not because it was out of reach but because the expected value did not justify the effort. That is a different selection mechanism from either obscurity or difficulty, and apple-picking does not contain it: a reach ceiling is about height, whereas this is about whose time is worth spending. It predicts AI’s comparative advantage lies in low-probability high-effort searches regardless of their depth, which would explain the hardened-code cyber finds as well [→ Mythos]. This reading is the log’s, not the author’s.
  • Its conclusion on demand. “In the short to medium term, this points to a world where AI models complement humans but do not replace them.”
  • Dates: published 2026-05-28, eight days after the disproof was announced.
  • Bears on: Q3 demand, Q4 expertise, Q7 incidence.
  • Unused: not yet cited in the argument. Its grind-not-height mechanism is a live alternative to the reach-ceiling story and belongs wherever the argument discusses why particular problems fell.
  • Links: Understanding AI
  • Status: verified as to the quotations, which were read from the article, retrieved 2026-07-26; journalism rather than research, so the mechanism claims are the author’s interpretation and carry no measurement behind them.

Algorithms

The first eight entries measure the rate itself: how fast the compute, time, or money needed to reach a fixed result falls. They sit here rather than in aggregate measures because algorithmic efficiency is this domain’s outcome variable, not a backdrop to it. Seven of the eight end before the agent era, so they establish the pre-AI rate that an AI contribution would have to beat, and none of them measures one; the exception is the inference-price series, which runs into the agent era but measures deployment cost rather than research output. Three of them — Bixby on solvers, Grace’s six-domain survey, and the SAT Museum — are also the only entries in this log measuring algorithmic progress in the classical sense of better exact algorithms, which is what the agent benchmarks further down do not do.

The entries after those eight are the agent benchmarks and demonstrations, then the field evidence on software work. The two entries auditing the optimization benchmarks stay here because that is what they audit; benchmark-validity work that cuts across domains is collected in Measurement and benchmark validity instead.

Sherry and Thompson: how fast do algorithms improve? (2021)

Independent (MIT CSAIL / MIT IDE). The widest survey of algorithmic progress there is: 113 algorithm families traced from the 1940s to 2019, scored by asymptotic worst-case complexity. It is the closest thing in the literature to a direct test of apple-picking, because it measures the distribution of improvement across problems rather than the average.

Two log-scale line charts. The upper panel plots relative performance from 1940 to 2020 for four algorithm families as step functions with sudden large jumps, against a smooth grey hardware improvement curve. The lower panel plots nearest-neighbour search against hardware by years since start, with the algorithmic gain increasing with problem size n.

Algorithmic improvement arrives in discrete jumps; hardware improvement is smooth. Panel (b) shows the same algorithmic jump being worth more at larger problem size.

Hardware improvement is the smooth grey staircase; each algorithm family is a flat line punctuated by one or two enormous discontinuities, and some families never move at all. The lower panel makes the second point: the same algorithmic discovery is worth about 2^4 at n = 100 and about 2^22 at n = 10^8, so whether an algorithm beats hardware depends on the problem size you ask about.

  • Enormous heterogeneity is the headline, not the average. “Analyzing data from 57 textbooks and more than 1137 research papers reveals enormous variation. Around half of all algorithm families experience little or no improvement. At the other extreme, 14% experience transformative improvements, radically changing how and where they can be used. Overall, we find that, for moderate-sized problems, 30%-43% of algorithmic families had improvements comparable or greater than those that users experienced from Moore’s Law and other hardware advances.”
  • The distribution is bimodal, which is the apple-picking shape. “The first cluster, representing just under half the families, shows little to no improvement even for large problem sizes.” And: “The second cluster of algorithms, consisting of 14% of the families, has yearly improvement rates greater than 1000% per year.” A mean improvement rate over these families would describe almost none of them.

Three stacked histograms of the percentage of algorithm families by average percentage improvement per year, with bins from 0-10% up to greater than 1000%, shaded to mark slower and faster than hardware. The mass at 0-10% falls from 64% to 45% as problem size grows while the greater-than-1000% bar stays at 14%.

The distribution of improvement rates across algorithm families, at problem sizes n = 1 thousand, 1 million, and 1 billion.
  • Whether algorithms beat hardware depends entirely on problem size. “For n = 1 thousand, only 18% of families had improvement rates faster than hardware, whereas 82% had slower rates. However, for n = 1 million and n = 1 billion, 30% and 43% improved faster than hardware. Correspondingly, the median algorithm family improved 6% per year for n = 1 thousand but 15% per year for n = 1 million and 28% per year for n = 1 billion. At a problem size of n = 1.06 trillion, the median algorithm improved faster than hardware performance.”
  • Improvements are rare events per family. “There are 113 algorithm families. On average, there are eight algorithms per family… there are 276 initial algorithms and subsequent improvements, an average of 1.44 improvements after the initial algorithm in each algorithm family.” Roughly one and a half improvements per family across eight decades is a very low arrival rate for a picked apple.
  • The authors’ summary, in their own words. “We find enormous heterogeneity in algorithmic progress, with nearly half of algorithm families experiencing virtually no progress, while 14% experienced improvements orders of magnitude larger than hardware improvement (including Moore’s law). Overall, we find that algorithmic progress for the median algorithm family increased substantially but by less than Moore’s law for moderate-sized problems and by more than Moore’s law for big data problems.”
  • The bias runs toward overstating progress, and they say so. “If inflation of leading constants is typical, it would mean that our results overestimate the scale of algorithm improvement.” They test it and find such cases are “the exception, rather than the rule,” but decline to extrapolate: “we cannot assume that this necessarily extrapolates to unmeasured algorithms since higher complexity may lead to both higher leading constants and a lower likelihood of quantifying them.”
  • What is excluded matters for reading it against this log. Only exact algorithms with exact solutions count.4 So most of what the agent benchmarks in this section actually reward — constant-factor speedups, better kernels, BLAS substitutions, parallelism — is by construction invisible here. The two literatures measure different things and neither substitutes for the other.
  • Why it is the sharpest conceptual evidence in the log. Apple-picking predicts that a fixed body of problems yields a few large gains and a long tail of nothing, and that the reachable set depends on where you look. This is that prediction, measured across eight decades, without any AI in it — which cuts both ways: it is the strongest confirmation of the shape and the strongest evidence that the shape is normal rather than AI-induced. That reading is the log’s, not the authors’.
  • Do not use MIT FutureTech’s summary of this paper. Its page reports “13%” and “30% to 45%”; the paper says 14% and 30%-43%.
  • Dates: data cover algorithms discovered from the 1940s to 2019, counting improvement years “since 1940”; early access on or about 2021-09-20 (Crossref record created that day, MIT News released the same day); print issue Proceedings of the IEEE 109(11), November 2021, pp. 1768–1777. No preprint exists. Read 2026-07-26.
  • Bears on: Q1 growth rate, Q5 returns, Q6 intertemporal, Q7 incidence.
  • Links: DOI 10.1109/JPROC.2021.3107219 · accepted manuscript, CC BY
  • Unused: not yet cited in the argument, and it should be — it is the base rate the whole apple-picking claim needs.
  • Status: verified — full text read from the CC BY accepted manuscript, whose abstract matches IEEE Xplore. Both figures reproduced from that PDF. IEEE Xplore itself is behind a bot check, so the day-level online date is inferred rather than read.

4 Sherry and Thompson, §III Methods: “This ‘exact algorithm, exact solution’ criterion also excludes, amongst others, algorithms where solutions, and even in theory, are imprecise (e.g., detect parts of an image that might be edges) and algorithms with precise definitions but where proposed answers are approximate. We also exclude quantum algorithms from our analysis since such hardware is not yet available.” And on what counts as an improvement: “We assess that an algorithm has improved if the work that needs to be done to complete it is reduced, asymptotically. This, for example, means that a parallel implementation of an algorithm that spreads the same amount of work across multiple processors or allows it to run on a GPU would not count toward our definition.”

Bixby: LP and mixed-integer programming solver speedups (2012)

Independent (author is Gurobi’s founder, so read the Gurobi figures as vendor-adjacent). The canonical measurement of algorithmic progress in optimization solvers, and methodologically the cleanest in this log: he removed hardware from the comparison physically rather than statistically.

Bar chart of version-to-version speedups for eleven consecutive CPLEX version pairs, with a line on a right-hand log axis showing cumulative speedup rising from about 3 to about 30,000.

Version-to-version and cumulative CPLEX speedups on the 1,892-model benchmark, CPLEX 1.2 through CPLEX 11.

The chart from Bixby’s paper. Bars are the version-to-version speedups on the left axis; the line is their cumulative product on the right-hand log axis, ending near 30,000× at CPLEX 11. Bixby picks out two bars. The first, CPLEX 2.1 to 3.0, is “an improvement factor of nearly 5.5,” which “corresponds to the maturity of the dual simplex algorithm.” The second, CPLEX 6.0 to 6.5, “occurred in 1998, a speedup exceeding a factor of 10.0.” He also notes one exception to the pattern of steady gains: every version “with the arguable exception of CPLEX 6.0, represented a significant improvement over the previous version.”

Log-scale scatter of cumulative speedup against year, rising in a near-straight line from 1 in the early 1990s to about 500,000 by 2012.

Cumulative machine-independent speedup on the same benchmark plotted against calendar year, reaching about 500,000×.

The same benchmark against calendar year rather than version pair, plotted by Hernandez and Brown and extended past Bixby’s 2007 test [→ Hernandez and Brown]. They read the slope as “a 2x speedup every 13 months,” and attribute the smoothness to aggregation: “The smooth progress is partially explained by the measure being an aggregation of many problems of varying difficulty.”

  • The method is the reason to trust it. “In late 2007, I undertook a massive computational test using the CPLEX codes that had been released over the years… From this extensive library, a test set of 1892 representative models was selected. Using these models, and using a bank of identical computing machines, I recompiled each of the corresponding twelve CPLEX released versions – from Version 1.2 (the first version having MIP) through CPLEX 11 – to run on the target machine.” Every version faces identical hardware, so the ratio is machine-independent by construction, not by regression.
  • MIP improved by a factor of over 29,000 from algorithms alone. “It was computed by multiplying the effects of the individual improvements, producing a projected, machine-independent improvement of a factor of over 29,000.”
  • LP: 3300× algorithmic against 1600× hardware, 1988 to 2004. “In a period of sixteen years, from 1988 to 2004, by at least some measure, the average speed of at least one LP code – independent of any machine effects – improved by a factor of roughly 3300, far in excess of the improvements in the speed of computing machines over that same period; moreover, combining the effects of the algorithms and the machines gives an improvement factor exceeding six orders of magnitude.”
  • The published table contains an arithmetic typo, in the original. The summary table gives “Algorithmic improvement (machine independent)… 3300×”, “Machine improvement: 1600×”, and “Total improvement (3300 · 2000): 5,280,000×”. The parenthetical says 2000 but the product is consistent with 1600 (3300 × 1600 = 5,280,000). Quote the row as printed and use 1600.
  • The LP figure is reported here, not derived here. Bixby writes only that “In [7] I reported in detail on the overall improvements in the CPLEX LP code from 1988 through 2002, and subsequently updated these results in 2004.” The LP methodology lives in his 2002 Operations Research paper; only the MIP method is documented in this one. So the two headline numbers do not have equal evidentiary standing.
  • His definition of “the algorithm” is generous, and he flags it. “Note that we have used here as our algorithm the best of barrier, primal, and dual. One can argue whether this is a legitimate approach, but it is the one that I have used. It means that, for each model in the test set, each of the three algorithms was run, and the solution time of the fastest of the three was taken as the solution time for the model.” Taking the per-instance minimum over three algorithms inflates the measured speedup relative to any single algorithm.
  • Progress stopped, which is the part usually dropped when this paper is cited. “Since 2004 there have been essentially no improvements in the standard LP algorithms, means that LP is threatening in the future to again become a significant bottleneck in our ability to solve real-world problems of interest.” An eight-year plateau after a 3300× run is the apple-picking shape in a single well-studied problem, stated by the person best placed to know.
  • Gains are lumpy, not smooth. “CPLEX 2.1 was approximately 3.1 times faster than CPLEX 1.2”; the CPLEX 3.0 jump was “nearly 5.5”, attributed to “the maturity of the dual simplex algorithm”; and “the second and by far the biggest improvement occurred in 1998, a speedup exceeding a factor of 10.0.” This is the same step-function pattern Sherry and Thompson chart [→ Sherry and Thompson], seen inside one product line.
  • Do not propagate the combined MIP figure. He reports 16.2× more from Gurobi 1.0 to 5.0 “on top of the factor of 29,000,” but never multiplies them out; any ~470,000× attributed to him is a secondary-source computation.
  • Dates: LP measurement spans 1988–2004; the MIP test was run in late 2007 over CPLEX versions released 1991–2007; the Gurobi extension covers 2009–2012. Published in Documenta Mathematica, Extra Volume: Optimization Stories (ISMP 2012), pp. 107–121; Crossref gives 2012-01-01 and the PDF’s embedded creation date is 2012-07-25, so cite the year only. Read 2026-07-26.
  • Bears on: Q1 growth rate, Q5 returns, Q6 intertemporal, Q7 incidence.
  • Links: DOI 10.4171/DMS/6/16 · open-access PDF
  • Unused: not yet cited in the argument. The plateau claim is the most citable part.
  • Status: verified — full text read from the open-access PDF. Bixby founded Gurobi, so the Gurobi-era figures are self-reported; the CPLEX-era figures predate the company and are documented in method. Bixby’s own version-speedup chart is reproduced from his paper; the calendar-year plot of the same benchmark is reproduced from Hernandez and Brown.

Grace: algorithmic progress in six domains (2013)

Independent (MIRI technical report). The earliest attempt to put algorithmic progress and hardware progress on the same scale across unrelated fields. Preliminary by its own description, but it is where the “algorithms are worth about as much as hardware” figure originates.

Bar chart of solve-time ratios for 189 SAT instances, rising from near zero on the left through 1.0 and truncated at 2 on the right.

Ratio of later-competition to earlier-competition solve time for each of 189 SAT instances, ordered from largest to smallest improvement.

The report’s Figure 1. Each bar is one of 189 SAT instances solved in two consecutive competitions, plotted as the ratio of the later solve time to the earlier one and ordered from largest improvement to smallest; the axis is truncated at 2, though Grace notes “a few problems took two to ten times longer in the second year.” Her reading of the distribution: “the distribution is almost uniform between zero and one, with a small fraction taking much longer than before. There is a flat spot: around six percent of problems’ times changed by less than one percent between years.” Bars above 1 are instances that got slower.

  • The headline, with its hedge attached. “We examine evidence of progress in six areas of algorithms research… Many of these areas appear to experience fast improvement, though the data are often noisy. For tasks in these areas, gains from algorithmic progress have been roughly fifty to one hundred percent as large as those from hardware progress. Improvements tend to be incremental, forming a relatively smooth curve on the scale of years.”
  • The six domains, from the report’s own section headings. Boolean satisfiability; game playing (chess and Go); factoring; physics simulations; operations research (linear and mixed-integer programming, and scheduling); machine learning. Secondary summaries disagree — AI Impacts splits chess from Go and drops physics simulations, MIRI’s announcement collapses operations research to mixed-integer programming — so use the report’s headings.
  • Per-domain rates. SAT solvers “5–15% per year, depending on the type of problem”; chess “around fifty Elo points per year over the last four decades”; Go “about one stone per year for the last three decades”; factoring “about 5.5 digits per year for the last two decades”; MIP algorithms “have roughly doubled in speed each year”; and machine learning has had “steeply diminishing progress in percentage accuracy over recent decades.”
  • The author’s own selection warning is the most valuable thing in the report. Grace states plainly that her sample is biased optimistic.5 That applies with full force to this log, which is built almost entirely out of benchmarks chosen because something interesting happened on them.
  • She flags her own retrospective-versus-prospective problem in the SAT data. “In recent Boolean satisfiability (SAT) competitions, SAT solver performance has increased 5–15% per year… However, these gains have been driven by widely varying improvements on particular problems. Retrospective surveys of SAT performance (on problems chosen after the fact) display significantly faster progress.” Same field, and the measured rate depends on when the problems were picked.
  • She is similarly careful about MIP, which is the fastest number in the report. MIP “is an important optimization problem, but one which has been called to attention after the fact due to performance improvements. Other optimization problems have had more inconsistent (and harder to determine) improvements.”
  • It is explicitly a first pass. “This has been a preliminary survey,” and the introduction calls it “a collection of first glances… This paper will neither analyze the data extensively nor attempt to point out all of its interesting implications.” There are no confidence intervals anywhere in it; the “low confidence” description that circulates comes from AI Impacts, not from Grace.
  • Her CPLEX test-set size differs from Bixby’s. Grace reports CPLEX “versions 1.2 to 11.0, released in 1991 to 2007” tested on 1,852 MIP problems, citing a 2010 Bixby talk; the 2012 paper says 1892 [→ Bixby]. Probably different presentations of the same experiment, but quote whichever document you are using.
  • Dates: released on or about 2013-08-03 (the LessWrong announcement is timestamped 2013-08-03 02:29:21Z and opens “Today MIRI released a new technical report”; the MIRI newsletter of 2013-08-13 also announces it), revised 2013-12-09, which is the only date printed inside the document. MIRI’s own announcement post carries no retrievable date stamp. Read 2026-07-26.
  • Bears on: Q1 growth rate, Q7 incidence, Q8 benchmarks.
  • Links: PDF · AI Impacts summary
  • Unused: not yet cited. The selection caveat is the part this log most needs to quote against itself.
  • Status: verified — full report read. Note that it is a 2013 technical report with no peer review and no journal version. Figure reproduced from the source.

5 Grace, §3.3 “On Selection”: “The most salient algorithmic problems might be those for which progress is particularly fast (or slow), so looking at algorithms that are being reported on might give us a biased impression of the overall rate of progress.” And: “In many cases, we should treat estimates as being optimistic rather than representative. We should rely more on assessments that are planned in advance of knowledge about performance. Competitions are better than retrospective analyses, and problems that were singled out early are better than problems that were selected after some progress.” And: “That which is easily measured may improve faster than more nebulous qualities, particularly if such measures are being used to guide progress. Yet we may care about the nebulous qualities. […] Thus, progress on well-defined metrics, such as most of what we will examine here, will tend to overestimate the progress we care about.”

Epoch AI: algorithmic progress in language models (2024)

Independent (research organization, peer-reviewed). Ho, Besiroglu, Erdil, Owen and co-authors fit augmented scaling laws to over 200 language model evaluations on WikiText and Penn Treebank spanning 2012–2023, estimating how fast the compute needed to reach a fixed performance threshold falls. (An earlier version of this entry credited the paper to Erdil and Besiroglu; Anson Ho is first author of a nine-author paper.) This is the best-measured algorithmic-efficiency curve in the algorithms domain, and the reference point for what a research-efficiency series looks like when someone builds one properly.

Stacked area chart of effective compute relative to 2014, with a large compute-scaling band and a smaller algorithmic-progress band.

Effective compute since 2014, decomposed into compute scaling and algorithmic progress.

Compute scaling contributed a factor of about \(1.7\times10^{7}\) against algorithmic progress at about \(2.2\times10^{4}\) — so even the best-measured algorithmic-efficiency curve is a small share of observed capability gains, and it is measured over a period ending before agents.

  • Compute to reach a fixed performance level halves roughly every 8 months. The abstract states it precisely: “using a dataset of over 200 language model evaluations on Wikitext and Penn Treebank spanning 2012-2023, we find that the compute required to reach a set performance threshold has halved approximately every 8 months, with a 95% confidence interval of around 5 to 14 months, substantially faster than hardware gains per Moore’s Law.”
  • The published confidence intervals disagree between versions. The arXiv abstract gives “a 95% confidence interval of around 5 to 14 months”; the NeurIPS 2024 version reports a 90% interval of about 2 to 22 months. Quote the median and treat the interval as wide. This log has not established which interval supersedes the other.
  • Most observed progress came from scale, not algorithms — and the authors say so themselves. “Despite the rapid pace of algorithmic progress and the development of new architectures such as the transformer, our analysis reveals that the increase in compute made an even larger contribution to overall performance improvements over this time period.” A Shapley decomposition attributes 60–95% of gains to compute and data and 5–40% to algorithmic innovation, with the algorithmic share falling as compute scaling accelerated after about 2018; those shares come from the Epoch summary page rather than the abstract.
  • Two named innovations are sized in time-equivalent units. The transformer is worth almost two years of algorithmic progress; Chinchilla scaling laws are worth 8 to 16 months. From the Epoch summary rather than a full read.
  • A direct check broadly agrees. Matching Megatron-LM or GPT-2 level performance required 5- to 100-fold less compute per year by 2023, implying a halving time the arXiv version puts at 6–15 months and the NeurIPS version at 11–17 months.
  • The authors’ own limitation. “Though limited by noisy benchmark data, our analysis quantifies the rapid progress in language modeling, shedding light on the relative contributions from compute and algorithms.” The hedge is the first clause.
  • Scope caveat: this is the pre-agent baseline, not a measurement of AI’s contribution. It covers pretraining only, ends in 2023, and uses perplexity benchmarks, so it describes the efficiency curve that AI research agents would have to bend, not any bending they have done. This log’s framing, not the paper’s.
  • Dates: evaluations span 2012–2023; arXiv 2024-03-09; NeurIPS version December 2024; retrieved 2026-07-26.
  • Bears on: Q1 growth rate, Q8 benchmarks.
  • Links: Epoch summary · arXiv · NeurIPS 2024 PDF
  • Status: verified-abstract — checked against the arXiv abstract, the Epoch summary page, and the NeurIPS abstract, retrieved 2026-07-26. The decomposition figure above is Epoch’s own chart from the summary page. The body figures above (Shapley shares, transformer and Chinchilla equivalents, direct-check halving times) come from those pages rather than a full read of the paper, and the interval discrepancy between versions is unresolved.

Hernandez and Brown: measuring the algorithmic efficiency of neural networks (2020)

Independent (OpenAI). The first paper to define algorithmic progress as compute-to-reach-a-fixed-past-capability, which is the measurement convention almost everything else in this section inherits.

Log-scale scatter of teraflop/s-days against year from 2012 to 2019; blue points mark the lowest-compute model at each date, from AlexNet down to EfficientNet-b0, with grey points above the frontier and a falling dashed trend line.

Training compute needed to reach AlexNet-level ImageNet performance, by year of model.

The paper’s Figure 3, titled “44x less compute required to get to AlexNet performance 7 years later.” Blue points are the lowest-compute model measured at any given time and grey points are all other models measured; the vertical axis is training compute in teraflop/s-days, held to a fixed accuracy target. The authors read the frontier as “an efficiency doubling time of 16 months.” They also flag that the target is sensitive to how far the original AlexNet was trained: at 62 epochs rather than 90, “we would have calculated the overall algorithmic efficiency gain as 30x rather than 44x.”

  • The framing, which is the contribution. “Algorithmic progress has traditionally been more difficult to quantify than compute and data. In this work, we argue that algorithmic progress has an aspect that is both straightforward to measure and interesting: reductions over time in the compute needed to reach past capabilities.” Fixing the capability and letting cost fall is what makes an efficiency series comparable to a cost curve.
  • The headline. “The number of floating-point operations required to train a classifier to AlexNet-level performance on ImageNet has decreased by a factor of 44x between 2012 and 2019. This corresponds to algorithmic efficiency doubling every 16 months over a period of 7 years.”
  • Algorithms beat hardware over the same window. “By contrast, Moore’s Law would only have yielded an 11x cost improvement.” The authors decline to treat these as rivals: “hardware and algorithmic efficiency gains multiply and can be on a similar scale over meaningful horizons, which suggests that a good model of AI progress should integrate measures from both.”
  • The headline figure is sensitive to a judgement call the authors flag themselves. Training AlexNet to convergence rather than to its near-final accuracy inflates the baseline.6 So the defensible range is 30–44×, and the 16-month doubling is the top of it.
  • Thin data, and they say so. “We only have a small number of algorithmic efficiency data points on a few tasks,” with the main result “primarily based on existing open source re-implementations of popular models.”
  • Why it is in this section. It is the earliest dated point in the efficiency-measurement series, it covers a period that ends before the agent era begins, and its 16-month doubling is almost exactly the transistor cost curve’s 17 months — a coincidence worth noticing but not over-reading.
  • Dates: data span 2012–2019; arXiv 2020-05-08; no later version. Retrieved and read 2026-07-26.
  • Bears on: Q1 growth rate, Q5 returns.
  • Links: arXiv 2005.04305
  • Unused: feeds the rate comparison below; not yet cited directly.
  • Status: verified — abstract and body read from the arXiv PDF. The OpenAI blog version is behind a bot check and was not usable. Figure 3 reproduced from the source.

6 Hernandez and Brown, §5: “It only took 62 of the 90 epochs for AlexNet to train to 78.8% top 5 accuracy on ImageNet (99.6% of the 79.1% final accuracy). So if the original AlexNet had only been trained for 62 epochs, we would have calculated the overall algorithmic efficiency gain as 30x rather than 44x. We don’t think it’s tractable to mitigate this confounder without adding a lot of complexity to explaining the measurement, but it seemed important to flag as a limitation of our approach.”

Erdil and Besiroglu: algorithmic progress in computer vision (2022)

Independent (Epoch AI). The ImageNet counterpart to the language-model estimate, by overlapping authors and a similar method, which is why the two should not be treated as independent confirmations.

Three log-log panels of compute against training-set size, with curves labelled 2012 through 2021 shifting downward and to the left over time.

Compute-data Pareto frontiers for reaching AlexNet, ResNeXt-101, and ViT-e performance, by year.

The paper’s Figure 1. Each curve traces the combinations of compute and training-set size sufficient to reach a fixed performance level in a given year, for three targets: AlexNet, ResNeXt-101, and ViT-e. The curves shift down and left from 2012 to 2021, which is algorithmic progress measured with the performance target held constant. The paper’s estimate is that compute-augmenting innovations “halve compute requirements every nine months (95% confidence interval: 4 to 25 months).”

  • Compute requirements halve about every nine months. “We estimate that compute-augmenting innovations halve compute requirements every nine months (95% confidence interval: 4 to 25 months).” The interval is very wide — wide enough to contain both “faster than anything in the historical record” and “slower than Moore’s law.”
  • Algorithms and compute contributed roughly equally. “Using Shapley values to attribute performance improvements, we find that algorithmic improvements have been roughly as important as the scaling of compute for progress computer vision.” Note this differs from the language-model finding, where compute dominated [→ Epoch on LMs] — same method, same group, opposite verdict on which input mattered more.
  • The split shifts over the period, which is the more useful fact. The equal-importance result comes “with the caveat that algorithmic progress was more crucial in early years, and compute scaling more important in later years.” A rising compute share as a field matures is what apple-picking predicts if the cheap algorithmic fruit goes first, though the paper draws no such inference.
  • Progress is compute-augmenting, not data-augmenting. “Algorithmic innovations mostly take the form of compute-augmenting algorithmic advances (which enable researchers to get better performance from less compute), not data-augmenting algorithmic advances.” The authors then withdraw part of the interpretation: “while our parameter estimates suggest algorithmic innovations are more effective at augmenting compute budgets relative to data budgets, these estimates alone do not necessarily imply this is also the primary mechanism through which algorithmic innovation improve performance.”
  • The model is noisy and the authors are direct about it. “The model has high uncertainty in some situations, both due to the parameter uncertainty and due to the noise that’s built into the model.”
  • It also supplies a secondhand summary of Grace (2013). The paper reports that Grace “investigates algorithmic progress in six domains (SAT solving, Game Playing, Factoring, Physics Simulations, Operations Research, and Machine Learning) and finds that, while the data is often noisy, gains from algorithmic progress have been roughly fifty to one hundred percent as large as those from hardware progress.” That is a secondary characterization; the primary report has not been read here.
  • Dates: arXiv 2022-12-10, current version v4 2023-08-24, read 2026-07-26. The estimate predates the agent era entirely.
  • Bears on: Q1 growth rate, Q5 returns, Q7 incidence.
  • Links: arXiv 2212.05153
  • Unused: feeds the rate comparison below; not yet cited directly.
  • Status: verified — abstract and body read from the ar5iv HTML. Figure reproduced from the source.

The SAT Museum: thirty years of solvers on one machine (2023)

Independent (academic: Biere, Fleury, Froleyks, and Heule). Historic SAT solvers recompiled and re-run on identical hardware, which makes the measured gain pure software — the same method Bixby used for CPLEX [→ Bixby], applied to an open competition series rather than one product line. It is the best pre-LLM algorithmic-progress baseline in the log for the shape of progress, and it settles a live dispute about whether the progress happened at all.

  • The design, which is why it can be trusted. “The virtual SAT Solver Museum is an effort towards preserving historical SAT solvers, by collecting and porting their source code to modern compilers and evaluating them on representative benchmark sets on the same hardware. This allows us to compare historic and modern solvers in the same environment. Our results clearly show a remarkable improvement of SAT solver performance in the last 30 years.”
  • Progress is mostly slow, with jumps every three to five years. “While in general we see a big improvement in solver performance in these 30 years across all considered benchmark sets, the yearly improvement is mostly rather slow, except for performance jumps in some years, which arguably happen with a frequency of 3 to 5 years.” The jumps are attributed to specific techniques: “one when preprocessing was introduced around 2006 and a second inconsistent one: 2019 on 2021/2022 benchmarks and 2016 on older benchmark sets,” the latter “attributed to a combination of local search (after 2019 in CaDiCaL and Kissat), rephasing (after 2016), and more aggressive bounded variable elimination.”
  • They are refuting a claim that gets made about this field. “It has been stated that ‘No major performance breakthrough [happened in SAT solving] in close to two decades’.” Their verdict: “Clearly, our presented results disprove the false view discussed in the introduction that there was no major progress in SAT solving in the last 20 years.”
  • They control for benchmark-selection bias explicitly, which almost nothing else here does. “We run on the same hardware (from 2016) all collected and patched solvers on six benchmark sets from SAT competitions spanning more than two decades, namely 2002, 2011, 2019, 2020, 2021, and 2022. We report results on each set separately, in order to address an argument brought forward by Laurent Simon at a recent POS workshop, that the benchmark selection method of more recent competitions might give a bias towards newer solvers … Our data on the SAT Competition 2011 benchmark set refutes this argument as it clearly shows the same solver progress which we observed in other years.”
  • Their own contamination caveat, which is the same problem the AI benchmarks have. “Regarding results we want to stress again that SAT solver developers train on previous competitions: The winner of the 2022 competition is based on the winner from 2021, had access and most likely has been trained on the 2021 benchmarks to find better heuristics than its predecessor from 2021, which has not been trained on them nor on more recent problems sets.”
  • Their own limitation. “one nagging remaining issue with our work is that we do not provide a deeper understanding about the differences between solvers and whether all implemented techniques are useful.”
  • No headline speedup factor exists, and this entry does not invent one. Results are reported as cumulative distribution functions rather than as a multiple, so any instances-solved count has to be read off the paper’s figures. The experimental setup, needed before reusing anything: “We ran all the benchmarks on 8-core Intel Xeon E5-2620 v4 CPUs running at 2.10 GHz (turbo-mode disabled) with a memory limit of 127 GB and a time limit of 5 000 seconds as in recent SAT competitions even though on slightly slower hardware.”
  • Dates: presented at the 14th International Workshop on Pragmatics of SAT, 2023; published in CEUR-WS Vol-3545, dated 2023. Precursors include a 2020 workshop lightning talk. No arXiv version. Read 2026-07-26.
  • Bears on: Q1 growth rate, Q4 expertise, Q6 intertemporal, Q7 incidence.
  • Unused: not yet cited in the argument. The three-to-five-year jump frequency is the most citable part, because it is a staircase measured in a domain with no AI in it — which is a problem for reading a staircase as an AI signature.
  • Links: PDF · project page and solvers
  • Status: verified — full text extracted from the PDF and all quotes read from it, retrieved 2026-07-26. A workshop paper rather than a refereed journal article, and the solved-instance counts are in figures rather than in text.

Epoch AI: LLM inference price declines (2025)

Independent (research organization). The deployment-cost counterpart to the pretraining-efficiency curve, and the one efficiency series in the log already denominated in money per unit of capability. Its spread is so wide that the entry is as much a warning about the measurement as a rate.

  • The rate is a range spanning two orders of magnitude, not a number. “The rate of decline varies dramatically depending on the performance milestone, ranging from 9x to 900x per year.” One anchored case: “The price to achieve GPT-4’s performance on a set of PhD-level science questions fell by 40x per year.”
  • The estimate depends heavily on the window chosen. “When we removed all model data before January 2024…the median rate increased from 50x per year to 200x per year.” A fourfold move in the median from a window choice is the reason no single figure from this entry should be quoted as the rate.
  • Epoch’s own hedge on persistence. “The fastest price drops in that range have occurred in the past year, so it’s less clear that those will persist.”
  • What it is and is not evidence about. Falling inference price is a joint product of hardware, algorithmic efficiency, competition, and margin decisions, so it is not a measurement of algorithmic progress and still less of AI’s contribution to research. It matters for Q5 and Q6 in a different way: it is the price of the input whose returns the argument is asking about, and it fell fast over exactly the period the cyber entries measure. This framing is the log’s.
  • Dates: published 2025-03-12; the underlying model-price data run to early 2025. Retrieved 2026-07-26.
  • Bears on: Q1 growth rate, Q5 returns, Q6 intertemporal.
  • Unused: not yet cited in the argument. It belongs in the Q6 discussion, since a falling input price is an independent reason to defer spend that has nothing to do with capability rising.
  • Links: Epoch data insight
  • Status: verified — quotes and publication date checked against the Epoch page, retrieved 2026-07-26. The underlying price database was not inspected.

AlphaEvolve (2025)

Vendor (Google DeepMind), paper + production claims. Evolutionary coding agent powered by Gemini models (Novikov et al. 2025).

  • Matrix multiplication. A 4×4 complex-valued matrix multiplication in 48 scalar multiplications, beating Strassen’s 49 for that setting — “the first improvement, after 56 years, over Strassen’s algorithm in this setting.”
  • Open problems. Matched state of the art on ~75% and improved it on ~20% of the 50-plus open problems to which it was applied; also a Verilog rewrite of a TPU matrix-multiplication circuit.
  • Dollar impact (May 2026 one-year update, self-reported). The 23% kernel speedup → ~1% of Gemini training time; a data-center scheduling heuristic recovers ~0.7% of fleet-wide compute, “in production for over a year,” valued at ~$500M/year. The 23% / 1% / 0.7% figures appear in both paper and blog; the $500M is press-only.
  • Dates: blog announcement 2025-05-14; arXiv 2025-06-16; one-year production update 2026-05-07.
  • Bears on: Q1 growth rate, Q2 autonomy, Q7 incidence.
  • Links: arXiv 2506.13131 · DeepMind blog · one-year update
  • Status: paper figures verified-abstract; production figures vendor.

TTT-Discover (2026)

Independent (academic, Jan 2026). Test-time reinforcement learning for discovery, using the open gpt-oss-120b model (Yuksekgonul et al. 2026).

Reward curves over test-time training steps comparing test-time training with best-of-N sampling and a no-test-time-training baseline.

Reward over test-time training steps, against best-of-N and no-adaptation ablations.

Best-of-N jumps early and then flattens; updating the policy at test time keeps a slope for longer. Here the harness rather than the model generation moved the frontier.

  • TriMul kernels. On the GPU MODE TriMul kernel (used in AlphaFold), discovered kernels reduced runtime by more than 15% relative to the best human submission on every tested GPU type. Largest on A100: 2,198μs versus the human best of 4,531μs (51.5% lower); H100: 1,161μs versus 1,371μs. Trained against H100 timing only, making the cross-hardware result notable.
  • Harness, not generation. The gain came from the test-time-training harness on an older open model, not from a new base-model generation.
  • Cost. Reports SOTA kernels for “a few hundred dollars” of test-time compute — the exact accounting (failed runs, evaluator cost, engineering time) still needs to be pinned down before publication.
  • Dates: arXiv 2026-01-22 (v2 2026-02-05).
  • Bears on: Q2 autonomy, Q4 expertise, Q6 intertemporal.
  • Links: arXiv 2601.16175 · project page
  • Status: kernel numbers verified against the project table and paper; cost accounting unverified.

modded-nanogpt speedrun (2024–2026)

Independent (public leaderboard). Public competition to minimize GPT-2 training time at fixed loss; a ledger of cumulative human and AI contributions (Jordan and contributors 2026).

Log-scale plot of minutes to reach the target loss falling from 45 minutes in May 2024 to 1.32 minutes in May 2026, with AI-set and human-set records marked.

Record training time against date, with records attributed to AI or human authors.

Derived from the records listed in this entry. The curve is a clean efficiency series with dated authorship, which almost nothing else here has — but the visible flattening after early 2026 is the more useful feature, since the AI-set records cluster in the flat region where each contributes about one percent.

  • AI-set records. Official verified records by AI-agent companies: Record 32 (hiverge.ai, 2.625 min, Sep 11 2025), Record 60 (Locus/Intology, 1.765 min, Jan 16 2026), Record 69 (Aster, 1.528 min, Feb 2 2026), Record 72 (Station, 1.496 min, Feb 10 2026). Baseline 45 min (llm.c, May 28 2024); current record 1.320 min (Record 84, May 21 2026) — a ~34× reduction. [verified against the repo README]
  • Depth decomposition. The deep, durable gains are human: Muon optimizer (~21%), U-Net skip connections (~8%), Paired Head Attention. The AI records are shallow-mechanical: hiverge.ai ~1.2%, Locus ~0.9% (an explicit “fused triton kernel” — kernel fusion, not a new idea).
  • Dates: baseline 2024-05-28; AI-set records 2025-09-11, 2026-01-16, 2026-02-02 and 2026-02-10; current record 2026-05-21; README read 2026-07-26.
  • Bears on: Q1 growth rate, Q2 autonomy, Q3 demand.
  • Links: GitHub README
  • Status: verified against the README.

Karpathy autoresearch (2026)

Independent (individual researcher, ~Mar 2026). Karpathy left an agent tuning nanochat for ~2 days.

  • Result. The agent “found ~20 changes that improved validation loss… all of them were additive and transferred to larger models” — roughly 20 retained edits out of ~700 tried, yielding ~11% improvement in “Time-to-GPT-2.”
  • Character. The improvements were things like adjusting AdamW constants and scalar multipliers — Karpathy: not “novel, ground-breaking ‘research’ (yet), but all the adjustments are ‘real’.” He describes the agent as “cagey and scared” on open-ended problems.
  • Dates: repository created 2026-03-06 and last pushed 2026-03-26; Fortune coverage 2026-03-17. The run itself is described as about two days somewhere inside that window.
  • Bears on: Q1 growth rate, Q4 expertise.
  • Links: GitHub · Fortune
  • Status: unverified in detail (runtime, retained-change count, and transfer claims still need the exact commit/README/transcript linked).

AlphaTensor and AlphaDev: pre-LLM algorithm discovery (2022–2023)

Vendor (Google DeepMind), both peer-reviewed in Nature. The reinforcement-learning predecessors to FunSearch and AlphaEvolve. They matter here as a control: they produced real algorithmic improvements with no language model involved, which bounds how much of the current results should be attributed to LLMs specifically.

  • AlphaTensor beat a fifty-year record in matrix multiplication. It found faster algorithms for multiplying small matrices, the first improvement on Strassen-era results in that setting. AlphaEvolve’s 4×4 complex-valued result is a later step in the same line [→ AlphaEvolve].
  • AlphaDev produced sorting routines now running trillions of times a day. Its discovered routines were merged into the LLVM standard C++ library, which is a rare case in this log of an AI-discovered artefact entering production infrastructure with a verifiable adoption path.
  • Both were narrow, and that is the point. Contemporary commentary noted AlphaTensor “could do matrix multiplication, but basically nothing else.” The move to program search was what made the method general. So the capability that changed between 2022 and 2023 was breadth, which is Benjamin Jones’s coverage parameter rather than depth [→ Jones].
  • They also anchor the counterfactual for cost. These systems found genuine improvements with RL search and no frontier-model inference bill, so a 2026 result obtained for six figures of tokens is not automatically evidence that model capability is what produced it.
  • Dates: AlphaTensor published in Nature 2022-10-05; AlphaDev published in Nature 2023-06-07.
  • Bears on: Q1 growth rate, Q2 autonomy, Q7 incidence.
  • Unused: not yet cited in the argument. The Algorithms section would be stronger for noting that the pre-LLM baseline already included production-deployed algorithmic discoveries.
  • Links: AlphaTensor, Nature · AlphaDev, Nature
  • Status: verified-abstract — publication dates and headline claims checked against the Nature article pages, retrieved 2026-07-26; vendor results, peer-reviewed.

Cui and co-authors: pooled coding-assistant RCTs (2025)

Independent (academic, three randomized trials). The largest randomized measurement of AI coding assistants, and the direct counterweight to METR’s developer RCT. Trials at Microsoft, Accenture, and an anonymized Fortune 100 firm, pooling 4,867 developers.

  • Access to an AI coding assistant raised completed tasks by about 26%. The pooled estimate across the three firms, with the cross-firm design making it the most replicated firm-level result available.
  • It points the opposite way from the METR trial, and the difference is the finding. METR measured a 19% slowdown on mature high-standard repositories with experienced maintainers [→ METR RCT]. Both are randomized; they differ in the starting point, the task population, and the quality bar. Read together they say the sign of AI’s productivity effect is set by the setting, not by the tool — which is the starting-point-dependence claim of Q7, established by randomization rather than by comparing benchmarks.
  • The outcome is task count, not task value. Completed tasks in a corporate workflow are not deep results, and nothing in the design distinguishes a task that mattered from one that did not. This is the same measurement problem the optimization benchmarks have, in a field setting.
  • Dates: NBER working paper 33777, issued 2025-05-09; trials run before that date.
  • Bears on: Q1 growth rate, Q3 demand, Q4 expertise, Q7 incidence.
  • Unused: not yet cited in the argument. Pairing it with the METR RCT would let the Q7 discussion rest on two randomized trials with opposite signs rather than on a cross-benchmark comparison.
  • Links: NBER w33777
  • Status: unverified — the pooled effect and sample size are taken from secondary summaries, retrieved 2026-07-26, and have not been checked against the working paper itself. Verify before use.

RE-Bench (2024)

Independent (METR). The most rigorous human-vs-agent calibration: 7 AI-R&D environments each ≈8 hours of expert work; 71 8-hour attempts by 61 human experts (Wijk et al. 2025).

Best-score-at-k curves against total time budget, showing agents ahead at short budgets and human experts overtaking at longer ones.

Best score at k against total time budget for human experts and agents.

The crossing point is the object of interest. Agents lead at two hours and humans lead by eight, so any claim about relative capability here is really a claim about the budget at which the comparison is made.

  • The crossover. At a 2-hour budget the best agents score ~4× human experts; humans “exceed the best agent scores when given 8 hours”; at 32 hours humans reach ~2× the top agent.
  • Dates: arXiv 2024-11-22 (v2 2025-05-27).
  • Bears on: Q3 demand, Q5 returns, Q8 benchmarks.
  • Links: arXiv 2411.15114 · METR report
  • Status: verified.

METR: RCT on experienced open-source developers (2025)

Independent (METR), randomized controlled trial. The only randomized measurement in this log of what AI does to real work on mature codebases. 16 experienced developers completed 246 real issues drawn from their own large repositories (averaging 22k+ stars and over 1M lines of code, with about 5 years of prior contribution each); each issue was randomly assigned to allow or disallow AI. Tools were the February–June 2025 frontier, primarily Cursor Pro with Claude 3.5/3.7 Sonnet.

Point estimates with confidence intervals showing economics and ML expert forecasts of about 40 percent speedup, developer estimates of about 20 percent speedup, and an observed 19 percent slowdown.

Expert forecasts, developer self-reports, and the observed effect on completion time.

Every prior points one way and the measurement points the other, including the developers’ own estimates made after they had finished. Whatever else the study establishes, it establishes that self-reported and forecast productivity are not usable proxies here.

  • Allowing AI increased completion time by 19%. The sign is the finding: on mature repositories with high quality standards, early-2025 AI tooling made experienced maintainers slower, not faster.
  • Everyone predicted the opposite, including afterwards. Developers forecast a 24% speedup beforehand and still believed they had been sped up by 20% after finishing. Economics experts forecast 39% and ML experts 38% shorter completion times. Self-reported and forecast productivity are therefore not usable proxies for measured productivity.
  • The mechanism is overhead, not incapacity. Time saved on initial code generation was offset by prompting, waiting on generations, and reviewing and correcting output; developers accepted AI output unmodified less than 44% of the time.
  • The authors tested the obvious confounds. Around twenty properties of the setting that could a priori explain the slowdown were collected and evaluated; the effect was robust across those analyses, though experimental artifacts are not entirely ruled out. The METR write-up says twenty properties and the arXiv version says twenty-one — an immaterial version discrepancy, noted for traceability.
  • Scope caveats that matter for this project. 16 developers, one snapshot of tooling now well out of date, and a deliberately high-standard setting. METR frames it as a snapshot of early-2025 capability in one relevant setting, not a general estimate. It says nothing about agentic 2026 tooling.
  • Why it is load-bearing here. It is the human-side counterpart to SWE-fficiency and GSO: the same high starting point that collapses agent gains on mature repos also produced a negative effect for humans using AI. That is the strongest available evidence that measured efficiency gains in already-optimized settings can be zero or negative while perceived gains stay large.
  • Dates: tasks run February–June 2025; METR write-up 2025-07-10; arXiv 2025-07-12 (v2 2025-07-25).
  • Bears on: Q1 growth rate, Q3 demand, Q4 expertise, Q7 incidence.
  • Links: METR write-up · paper PDF · arXiv 2507.09089 · Ars Technica
  • Status: verified-abstract — headline effect, forecast gaps, expert forecasts, acceptance rate, and setting verified against the METR write-up and the paper abstract, retrieved 2026-07-26. The per-condition task split (136 AI-allowed, 110 AI-disallowed) and the confound analysis come from the arXiv PDF’s front matter rather than a full read.

AlgoTune (2025)

Independent (academic benchmark). 154 numerical-code optimization tasks with a hard $1-per-task LLM budget (Press et al. 2025).

  • Surface-level gains. “AlgoTuner achieves an average 1.72x speedup against our reference solvers, which use libraries such as SciPy, sk-learn and CVXPY. However, we find that current models fail to discover algorithmic innovations, instead preferring surface-level optimizations.” Canonical example: the communicability task gets >142× purely by “using BLAS operations instead of pure Python” — a library substitution, not a better algorithm.
  • The benchmark was built precisely to escape already-solved problems. “Evaluations have thus far focused on models’ performance on tasks that humans have previously solved, including in programming and mathematics. We therefore propose testing models’ ability to design and implement algorithms in an open-ended benchmark.” The tasks were chosen to give algorithmic invention a chance to appear, which is why its absence is informative rather than an artifact of the task set.
  • What the loop does, and on how small a budget. “AlgoTuner uses a simple, budgeted loop that edits code, compiles and runs it, profiles performance, verifies correctness on tests, and selects the fastest valid version,” under a hard $1-per-task LLM cap. That is a statement about cheap search, not about the frontier of spend — relevant when reading it against Q5.
  • Dates: arXiv 2025-07-19 (v4 2025-10-24).
  • Bears on: Q1 growth rate, Q7 incidence, Q8 benchmarks.
  • Links: arXiv 2507.15887
  • Status: verified-abstract.

MLGym (2025)

Independent (academic benchmark). ML research tasks for agents (Nathani et al. 2025).

  • Improvement without invention. Frontier models “can improve on the given baselines, usually by finding better hyperparameters, but do not generate novel hypotheses, algorithms, architectures, or substantial improvements.” The list of four things they fail to do is the useful part.
  • The tasks were designed to demand more than tuning. The benchmark holds “13 diverse and open-ended AI research tasks,” and solving them is meant to require “generating new ideas and hypotheses, creating and processing data, implementing ML methods, training models, running experiments, analyzing the results, and iterating through this process.” The gap between that intent and the hyperparameter search actually observed is the finding.
  • Models evaluated, which dates the claim. “Claude-3.5-Sonnet, Llama-3.1 405B, GPT-4o, o1-preview, and Gemini-1.5 Pro” — a February 2025 frontier, now well behind. Read the negative result as a datum about that generation, not a standing limit. This caveat is the log’s, not the paper’s.
  • Dates: arXiv 2025-02-20 — the oldest agent benchmark in this log.
  • Bears on: Q1 growth rate.
  • Links: arXiv 2502.14499
  • Status: verified-abstract.

SWE-fficiency (2025)

Independent (academic benchmark). Performance-optimization tasks in mature repositories (Ma et al. 2025).

Three panels plotting pre-edit runtime, gold speedup factor, and gold patch size against bins of achieved speedup ratio, for four models.

Agent speedup against how much headroom the original code had.

In the left two panels, agents do well only where the original code was slow and the expert patch achieved a large speedup, and collapse where the target was already near-optimal. The effect holds across four different models, so it is a property of the target rather than of the agent.

  • Agents reach a small fraction of expert speedup, and the figure recorded here was wrong. This entry quoted “less than 0.15× the expert speedup on average.” The v3 abstract says “on average, agents achieve less than 0.23x the expert speedup.” Use 0.23×. The direction of the finding is unchanged, but the earlier figure overstated the gap and was presented as a quotation.
  • The task and its scale. “498 tasks across nine widely used data-science, machine-learning, and HPC repositories (e.g., numpy, pandas, scipy): given a complete codebase and a slow workload, an agent must investigate code semantics, localize bottlenecks and relevant tests, and produce a patch that matches or exceeds expert speedup while passing the same unit tests.”
  • Why the benchmark exists, in the authors’ words. “Most benchmarks emphasize what to fix rather than how to fix code.” A how-to-fix benchmark on mature repositories is close to the setting where apple-picking predicts agents should do worst, which is why this entry carries weight for Q7.
  • The failure is localization, not editing. “Agents struggle in localizing optimization opportunities.” Finding what to improve in already-good code is the depletion story rather than a ceiling on writing code — and it matches the figure above, where agents succeed mainly where headroom was large.
  • Dates: arXiv 2025-11-08 (v3 2026-06-27).
  • Bears on: Q7 incidence, Q8 benchmarks.
  • Links: arXiv 2511.06090
  • Status: verified-abstract for the headline figure; the speedup-by-baseline figure above was read from the v3 HTML on arXiv, retrieved 2026-07-26, and is the strongest single piece of Q7 evidence in the log.

GSO (2025)

Independent (academic benchmark). Repository-level performance optimization (Shetty et al. 2025).

  • Leading agents fail almost completely. “Our quantitative evaluation reveals that leading SWE-Agents struggle significantly, achieving less than 5% success rate, with limited improvements even with inference-time scaling.” The second clause bears on Q5: more inference compute did not rescue it.
  • Construction. An automated pipeline analyzes “repository commit histories to identify 102 challenging optimization tasks across 10 codebases, spanning diverse domains and programming languages,” scoring the agent “against the expert developer optimization” — so the target is by construction a change a human actually made.
  • The named failure modes are diagnostic. “Difficulties with low-level languages, practicing lazy optimization strategies, and challenges in accurately localizing bottlenecks.” Lazy optimization and poor localization are the same pattern SWE-fficiency reports [→ SWE-fficiency], from an independent team and a different task set.
  • Dates: arXiv 2025-05-29 (v3 2025-10-24).
  • Bears on: Q7 incidence, Q8 benchmarks.
  • Links: arXiv 2505.23671
  • Status: verified-abstract.

PERFOPT-Bench relay pilot (2026)

Independent (academic, exploratory). Tests whether apparent within-run plateaus are real or artifacts of context management.

  • Relay result. Restarting an agent from an externalized optimization summary recovered additional headroom after the first session stalled — some apparent plateau is a context-management failure rather than exhaustion of reachable improvements. The paper labels this “an exploratory relay p[ilot]” and this log treats it as suggestive only.
  • Scaffold effects. “Optimization performance is workload-dependent rather than determined by model identity alone: no single stack dominates, and changing the agent framework can materially change the same LLM’s per-task speedup profile.” Evaluated over “7 agent stacks with different LLMs and agent frameworks on 7 long-horizon optimization tasks.”
  • Raw speedup is not a safe score, and they say so. “We further find that raw speedup is unsafe as a benchmark score, since some large gains arise from benchmark-specific shortcut exploitation.” This is the same warning the reliability audit reaches by a different route [→ benchmark reliability], and it is the sharpest Q8 evidence in the log.
  • What the tasks are. “Each task provides a correct but deliberately suboptimal codebase and asks the agent to improve a target performance metric; scoring requires hidden correctness tests, verified-speedup measurement, and trajectory-level audit.” Deliberately suboptimal is important: this is the opposite of the mature-repository setting, so it does not test depletion.
  • Caveat. The pilot covers only seven deliberately suboptimal tasks — it weakens a literal wall claim without overturning RE-Bench’s crossover.
  • Dates: arXiv 2026-07-08 — the newest entry in this log.
  • Bears on: Q5 returns, Q6 intertemporal, Q8 benchmarks.
  • Links: arXiv 2607.07744
  • Status: verified-abstract.

Performance-benchmark reliability audit (2026)

Independent (academic). “Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?” — replayed 740 reference patches across four machine types.

  • What is being audited, and why. “Their leaderboard scores are increasingly used as evidence of coding-agent progress, but those scores can conflate runtime instability, benchmark-specific scoring rules, and how many tasks are already solved by at least one public submission.”
  • Validity rates. After replaying “the official reference patches for 740 code optimization tasks across four common types of Google Cloud machines”: “most benchmark tasks can be replayed, but their reference patches satisfy the original benchmark validity rules in every cross-machine replay for only 39/102 GSO tasks, 11/140 SWE-Perf tasks, and 411/498 SWE-fficiency tasks; SWE-Perf is especially fragile because many reference patches produce close-to-zero runtime changes.” The qualifier “in every cross-machine replay” is doing real work and should not be dropped.
  • Scoring-rule sensitivity. “Public submission rankings depend strongly on the benchmark scoring rule. Among eight public submissions shared by GSO and SWE-fficiency, the official rankings disagree on 9 of 28 pairwise submission comparisons.” SWE-fficiency’s ten highest-weight tasks jointly receive 58.5–82.8% of total score weight.
  • Upshot. Absolute multipliers and rankings are machine- and scoring-rule-dependent. The broad result that agents struggle on repository-level optimization survives — it is reported by GSO and SWE-fficiency independently — but exact cross-benchmark comparisons should not bear weight. This reading is the log’s; the paper audits rather than adjudicates.
  • Dates: arXiv 2026-07-01 (v2 2026-07-16).
  • Bears on: Q8 benchmarks.
  • Links: arXiv 2607.01211
  • Status: verified-abstract.

SWE-bench: the benchmark that dated the floor at 1.96% (2023)

Independent (academic, Princeton, with the Verified subset produced by OpenAI). The default measure of agentic software engineering, and the origin of the number most lab announcements now quote. Its value here is its launch figure: the frontier could barely do the task at all in late 2023, which fixes the base of the curve everything else in this section sits on.

  • Scale and construction. “we introduce SWE-bench, an evaluation framework consisting of \(2,294\) software engineering problems drawn from real GitHub issues and corresponding pull requests across \(12\) popular Python repositories.”
  • The launch-date capability, which is the base of the curve. “Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere \(1.96\)% of the issues.”
  • What the task demands, in the authors’ words. “Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks.”
  • The Verified subset is a different object and should be named as such. The benchmark site describes it as “A human-filtered subset of 500 instances from SWE-bench,” where “Human annotators reviewed each instance to ensure the problem descriptions are clear, the test patches are correct, and the tasks are solvable given the available information.” The fraction of original instances judged underspecified or unfairly tested is not verified here, so scores on the two variants should not be joined into one series.
  • Dates: arXiv 2023-10-10 (v2 2024-04-05, v3 2024-11-11); the Verified subset was announced by OpenAI in August 2024, a date this log has not pinned to a primary page.
  • Bears on: Q1 growth rate, Q2 autonomy, Q8 benchmarks.
  • Unused: not yet cited in the argument. The 1.96% is the most useful thing in it — an efficiency series needs a base, and this is the only dated near-zero in the algorithms domain.
  • Links: arXiv 2310.06770 · Verified subset
  • Status: verified — abstract and version history read on the arXiv abstract page, and the Verified-subset language on the benchmark site, retrieved 2026-07-26. OpenAI’s own announcement returns an error to automated fetching and was not read; read the contamination critique alongside this entry [→ SWE-bench illusion].

MLE-bench: agents on Kaggle competitions against public leaderboards (2024)

Vendor (OpenAI), externally scored. An ML-engineering benchmark whose human baseline comes from Kaggle’s public leaderboards rather than from the authors, which makes the denominator unusually interpretable for a research-automation claim.

  • The task set and where the human baseline comes from. “we curate 75 ML engineering-related competitions from Kaggle, creating a diverse set of challenging tasks that test real-world ML engineering skills such as training models, preparing datasets, and running experiments. We establish human baselines for each competition using Kaggle’s publicly available leaderboards.”
  • The headline, with its full denominator and its threshold. The best setup “achieves at least the level of a Kaggle bronze medal in 16.9% of competitions.” Bronze is the threshold, and the figure is a share of competitions rather than a rank.
  • The authors raise contamination themselves. “In addition to our main results, we investigate various forms of resource scaling for AI agents and the impact of contamination from pre-training.”
  • Dates: arXiv 2024-10-09 (six later versions through 2025-02-26, the ICLR version).
  • Bears on: Q2 autonomy, Q4 expertise, Q5 returns, Q8 benchmarks.
  • Unused: not yet cited in the argument. It is the closest thing in the log to a dated, denominated measurement of end-to-end ML engineering, which is the activity the Algorithms section’s unobserved internal curve is about.
  • Links: arXiv 2410.07095
  • Status: verified-abstract — the quotes and the 16.9% checked against the arXiv abstract, retrieved 2026-07-26. Produced by a lab evaluating its own models among others.

PaperBench: replicating ICML papers from scratch (2025)

Vendor (OpenAI), with rubrics co-developed with the papers’ authors. The closest measurement in the log of end-to-end research execution, and it reports a human comparison that goes the other way from most benchmark headlines.

  • The task and its construction. “Agents must replicate 20 ICML 2024 Spotlight and Oral papers from scratch, including understanding paper contributions, developing a codebase, and successfully executing experiments.” And: “In total, PaperBench contains 8,316 individually gradable tasks. Rubrics are co-developed with the author(s) of each ICML paper for accuracy and realism.”
  • The headline, with its denominator. “the best-performing tested agent, Claude 3.5 Sonnet (New) with open-source scaffolding, achieves an average replication score of 21.0%.”
  • The human comparison, stated by the authors as a negative result. “Finally, we recruit top ML PhDs to attempt a subset of PaperBench, finding that models do not yet outperform the human baseline.”
  • Why replication is the right floor to measure. Reproducing a published result is strictly easier than producing a new one, because the target is known to exist and to be reachable. A 21% score on the easier task bounds what the same systems could be doing on the harder one, which is the inference CORE-Bench also invites [→ CORE-Bench]. This is the log’s reading.
  • Dates: arXiv 2025-04-02 (v3 2025-04-07).
  • Bears on: Q2 autonomy, Q3 demand, Q4 expertise, Q8 benchmarks.
  • Unused: not yet cited in the argument. It is direct evidence on autonomy in ML research, where the argument currently reasons from kernel records and leaderboard entries.
  • Links: arXiv 2504.01848
  • Status: verified-abstract — all figures and quotes checked against the arXiv abstract of v3, retrieved 2026-07-26.

CORE-Bench: computational reproducibility as the floor task (2024)

Independent (academic, Princeton). Reproducing published results from their own released code and data. The lowest bar in the log for a research task, which is what makes its failure rate informative.

  • Task set with denominators. “We introduce CORE-Bench (Computational Reproducibility Agent Benchmark), a benchmark consisting of 270 tasks based on 90 scientific papers across three disciplines (computer science, social science, and medicine). Tasks in CORE-Bench consist of three difficulty levels and include both language-only and vision-language tasks.”
  • The headline, with the authors’ own gloss. “The best agent achieved an accuracy of 21% on the hardest task, showing the vast scope for improvement in automating routine scientific tasks.”
  • The authors’ statement of why this is the prerequisite. “Having agents that can reproduce existing work is a necessary step towards building agents that can conduct novel research and could verify and improve the performance of other research agents.”
  • Dates: arXiv 2024-09-17 (v2 2026-06-22).
  • Bears on: Q2 autonomy, Q4 expertise, Q8 benchmarks.
  • Unused: not yet cited in the argument. Together with PaperBench it establishes that the easiest research task in the log is failed about four times in five, which is a useful constraint on autonomy claims elsewhere.
  • Links: arXiv 2409.11363
  • Status: verified-abstract — counts and quotes checked against the arXiv abstract, retrieved 2026-07-26. The 21% is from the version as read; the v2 revision was not compared against v1, so the figure may have moved.

The AI Scientist: automated papers at under fifteen dollars each (2024)

Vendor (Sakana AI, with academic co-authors). The first widely-cited end-to-end automated-research claim in ML. The cost figure is the load-bearing number, and the evaluator being the authors’ own is the load-bearing caveat.

  • Cost per unit of output. “Each idea is implemented and developed into a full paper at a cost of less than $15 per paper.”
  • The breadth claim, with its denominator. “We demonstrate its versatility by applying it to three distinct subfields of machine learning: diffusion modeling, transformer-based language modeling, and learning dynamics.”
  • The quality claim, and the fact that its judge is theirs. “To evaluate the generated papers, we design and validate an automated reviewer, which we show achieves near-human performance in evaluating paper scores. The AI Scientist can produce papers that exceed the acceptance threshold at a top machine learning conference as judged by our automated reviewer.” The last five words are the whole qualification, and they are usually dropped when this result is passed along.
  • The loop the authors describe. “In principle, this process can be repeated to iteratively develop ideas in an open-ended fashion, acting like the human scientific community.”
  • Dates: arXiv 2024-08-12 (v3 2024-09-01).
  • Bears on: Q2 autonomy, Q4 expertise, Q5 returns, Q8 benchmarks.
  • Unused: not yet cited in the argument. Its use is as the cheapest claimed research output in the log by two orders of magnitude, which makes it the natural test case for whether cost per result measures anything about value.
  • Links: arXiv 2408.06292
  • Status: verified-abstract — all quotes checked against the arXiv abstract of v3, retrieved 2026-07-26; vendor, and the quality evaluation is the vendor’s own automated reviewer rather than human review.

The AI Scientist-v2: one autonomous manuscript through real peer review (2025)

Vendor (Sakana AI). The successor, and the only claim in the log of an autonomously generated manuscript clearing an actual human review process. The denominator is the entire story.

  • The claim, with its denominator, in the authors’ words. “We evaluated The AI Scientist-v2 by submitting three fully autonomous manuscripts to a peer-reviewed ICLR workshop. Notably, one manuscript achieved high enough scores to exceed the average human acceptance threshold, marking the first instance of a fully AI-generated paper successfully navigating a peer review.”
  • What changed from v1, which bears on generality. “The AI Scientist-v2 eliminates the reliance on human-authored code templates, generalizes effectively across diverse machine learning domains, and leverages a novel progressive agentic tree-search methodology managed by a dedicated experiment manager agent.”
  • The authors place the achievement at workshop level themselves. The paper’s own subtitle is “Workshop-Level Automated Scientific Discovery via Agentic Tree Search.”
  • One of three, at a workshop, above an average threshold. Every qualifier there is doing work, and the entry records them together because the claim is usually compressed to “AI paper passes peer review.” A workshop acceptance threshold is not a conference one, and one acceptance from three submissions is a rate with a denominator of three.
  • Dates: arXiv 2025-04-10.
  • Bears on: Q1 growth rate, Q2 autonomy, Q4 expertise, Q8 benchmarks.
  • Unused: not yet cited in the argument. It is the strongest autonomy claim available in the algorithms domain and the one whose denominator most repays stating.
  • Links: arXiv 2504.08066
  • Status: verified-abstract — quotes checked against the arXiv abstract, retrieved 2026-07-26; vendor. The workshop, the reviews, and the scores were not independently examined.

Peng and co-authors: the Copilot randomized trial (2023)

Vendor-affiliated (Microsoft and GitHub authors, with an academic co-author), randomized controlled trial. The trial that set expectations for AI coding productivity, and the natural foil for METR’s negative result two years later [→ METR RCT]. Its narrowness is the point, and the authors state it themselves.

  • The headline, with the confidence interval that is usually dropped. “The performance difference between treated and control groups are statistically and practically significant: the treated group completed the task 55.8% faster (95% confidence interval: 21-89%).” An interval from 21% to 89% is consistent with a modest gain and with a near-doubling.
  • The task, which is where the external-validity problem lives. “Recruited software developers were asked to implement an HTTP server in JavaScript as quickly as possible.”
  • The effective sample is much smaller than the headline suggests. “A total of 166 offers were sent during the experiment, and 95 were accepted. The 95 developers were randomly assigned into control and treated groups, with 45 in the treated group and 50 in control. Thirty-five developers from both the treated and control groups completed the task and survey.” So the completion-time estimate rests on about 35 per arm, conditional on completing.
  • Success rate was a null. “We also find that the treated group’s success rate is 7 percentage points higher than the control group, but the estimate is not statistically significant, with a 95% confidence interval of [-0.11, 0.25].”
  • The authors’ own scope limits, quoted because they are more restrictive than the citation practice suggests. “This study examines a standardized programming task in an experiment to obtain a precise measure of productivity, instead of a task where developers collaborate on large projects in professional proprietary and/or open-source settings. Productivity benefits may vary across specific tasks and programming languages, so more research is needed to understand how our results generalizes to other tasks. Finally, this study does not examine the effects of AI on code quality.”
  • The expertise interaction runs against experience. “The results show that less experienced developers (years of professional coding), developers with heavy coding load (hours of coding per day), and older developers (developers aged between 25 and 44) benefit more from Copilot.”
  • Subjects underestimated their own gain, which is the opposite of METR’s finding. “both treated and control groups estimated a 35% increase in productivity, which is an underestimation compared with the 55.8% increase in their revealed productivity.” Read against METR, where developers believed they had been sped up 20% while being slowed 19%, the pair says self-report is unreliable in both directions rather than biased one way. This comparison is the log’s.
  • Dates: arXiv 2023-02-13; the tooling is the early-2023 non-agentic frontier.
  • Bears on: Q1 growth rate, Q3 demand, Q4 expertise, Q7 incidence, Q8 benchmarks.
  • Unused: not yet cited in the argument. Pairing it with the METR RCT and the pooled trials would give the Q7 discussion three randomized estimates whose signs differ by setting [→ METR RCT, pooled RCTs].
  • Links: arXiv 2302.06590
  • Status: verified — all quotes, the confidence intervals, and the sample accounting read from the arXiv PDF, retrieved 2026-07-26. Authors are at the firms selling the tool, and code quality was not measured.

Hoffmann and co-authors: Copilot shifts what developers do (2024)

Vendor-affiliated (academic authors using GitHub’s internal data), regression discontinuity. Exploits an eligibility threshold for free Copilot access on GitHub. The most useful non-experimental evidence in the log on task composition rather than speed, which is the object the task-replacement theory is about.

  • Design and scale. “We start with a panel of 187,489 distinct developers observed weekly from July 2022 through July 2024, which results in millions of developer-week observations for Copilot usage and activity levels in public GitHub repositories.”
  • The composition shift, with both absolute and relative magnitudes. “We find that coding activities as a percentage of all activity increase by 5.4 percentage points (12.37% relative to the baseline) while project management as a percentage of all activity drops by 10 percentage points (24.93% relative to the baseline).”
  • The two mechanisms, both directly relevant to this project. “We identify two underlying mechanisms driving this shift - an increase in autonomous rather than collaborative work, and an increase in exploration activities rather than exploitation. The main effects are greater for individuals with relatively lower ability.”
  • The ability interaction, restated. “We further find that the programming generative AI Copilot shifts the task allocation of developers with lower ability more than those with higher ability.”
  • The scope limit is that “lower ability” is within an already-selected population. The discontinuity is on an internal top-developer ranking, so the contrast is among top developers rather than across the whole distribution. The authors’ own framing: “Within the data set of top developers, we find that those who receive free access to Copilot…”.
  • Dates: HBS working paper 25-021, version dated 2024-10-27; the panel runs July 2022 to July 2024.
  • Bears on: Q2 autonomy, Q3 demand, Q4 expertise, Q6 intertemporal — exploration against exploitation is the intertemporal margin in a different guise.
  • Unused: not yet cited in the argument. It measures a shift from coordination to coding, which is the reallocation the Cyber section infers from curl and ARTEMIS without measuring it.
  • Links: working paper PDF
  • Status: verified — quotes read from the working-paper PDF, retrieved 2026-07-26. The regression-discontinuity diagnostics were not checked; the authors assert robustness to alternative estimators. Uses proprietary data from the firm selling the tool, with the firm’s employees as co-authors.

Song and co-authors: Copilot raises contributions and coordination cost together (2024)

Independent (academic authors, using GitHub Copilot usage data). Decomposes the project-level effect of Copilot on open-source repositories into a participation margin, an individual-productivity margin, and an offsetting coordination cost. The decomposition is the reason it is here: almost nothing else in the log reports an offsetting cost at all.

  • The headline with its decomposition. “we find that Copilot use increases project-level code contributions by 5.9%. This gain is driven by a 3.4% rise in developer coding participation and a 2.1% increase in individual productivity.”
  • The offsetting cost, in the authors’ words. “However, Copilot use also leads to an increase in coordination time by 8% due to more code discussions. This reveals an important tradeoff: While AI expands who can contribute and how much they contribute, it slows coordination in collective development efforts.”
  • The net sign, with its outcome variable named. “the combined effect of these two competing forces remains positive, indicating a net gain in overall project-level timely merge of code contributions from using AI pair programmers.”
  • Incidence across the skill distribution. “Peripheral developers show relatively smaller increases in project-level code contributions and experience larger increases in coordination time than core developers.”
  • Dates: arXiv 2024-10-02 (v3 2026-05-14).
  • Bears on: Q1 growth rate, Q3 demand, Q4 expertise, Q7 incidence.
  • Unused: not yet cited in the argument. The 8% coordination cost is the closest thing the log has to a measured version of the validation-bottleneck claim in software work [→ validation bottleneck].
  • Links: arXiv 2410.02091
  • Status: verified-abstract — all four quotes and the version history checked against the arXiv abstract, retrieved 2026-07-26. Observational rather than randomized; the identification strategy was not examined.

DORA: throughput up, delivery stability still down (2025)

Vendor-published survey (Google Cloud and the DORA research program). A repeated cross-section of software delivery practice. Its value is that one sign flipped between waves and the other did not, which is more informative than either wave alone.

  • The sign that changed, in DORA’s words. “Unlike last year, we observe a positive relationship between AI adoption on both software delivery throughput and product performance.”
  • The sign that did not. “AI adoption does continue to have a negative relationship with software delivery stability.”
  • Adoption and self-reported productivity, with the self-report visible in the wording. “90% of survey respondents report using AI at work” and “More than 80% believe it has increased their productivity.”
  • The denominator. “survey responses from nearly 5,000 technology professionals from around the world” plus “over 100 hours of qualitative data.”
  • How much weight it can carry. These are correlations in a self-selected survey with self-reported adoption and self-reported productivity, and the log’s own evidence ranking puts that near the bottom [→ METR RCT for why]. The persistence of the stability finding across two waves with an opposite-signed throughput finding is what makes it worth recording. This assessment is the log’s.
  • Dates: report announced 2025-09-23; the prior wave is the 2024 report.
  • Bears on: Q1 growth rate, Q3 demand, Q8 benchmarks.
  • Unused: not yet cited in the argument. Its use is as the quality-adjusted counterpart to throughput measures, which every benchmark in this section ignores.
  • Links: announcement · report landing page
  • Status: verified — quotes and the date checked against the announcement page, retrieved 2026-07-26; the full report PDF was not opened. Survey-based, self-selected, and published by a vendor.

GitClear: code duplication up, refactoring down (2025)

Vendor (GitClear, which sells code-quality analytics). The only source in the log with a year-by-year panel of what kind of change developers commit. Treat the framing as interested and the small denominators as limiting, but the composition series has no substitute here.

  • Dataset. “211 million changed lines of code, authored between January 2020 and December 2024” from “repos owned by Google, Microsoft, Meta, and enterprise C-Corps.”
  • The composition panel, read from the report’s own table rather than from a sentence. Lines classified “Moved” fall from 24.1% in 2020 to 9.5% in 2024, a stated year-on-year change of −39.9%; “Copy/pasted” rises from 8.3% to 12.3%; churn rises from 3.1% to 5.7%. Moved code is the signature of refactoring, so a collapse in that share alongside a rise in copy-paste is a composition shift away from consolidation.
  • They score their own prior forecast and report missing it. “The actual breakdown of code lines committed during 2024 was substantially worse than our projection.”
  • The duplicate-block result, with its much smaller denominator stated. From the duplicate-block table: 56,495 commits scanned in 2024, of which 3,764 contained a duplicate block, against 0.45% of 40,010 commits scanned in 2022. Their reading: “2024 was without precedent in the likelihood that a commit would contain a duplicated code block. The prevalence of duplicate blocks in 2024 was observed to be approximately 10x higher than it had been two years prior.”
  • The attribution gap, which is the reason this cannot support a causal claim. Nothing in the data identifies which lines an assistant wrote. The report’s mechanism claim is inference from use: “It’s readily apparent, from using these tools, that many of the suggested code blocks have their origins in existing code.” So the entry establishes a composition change over a period when AI adoption rose, not that AI caused it.
  • Dates: report version dated 2025-02-14; covers code authored January 2020 to December 2024. A 2026 successor exists and was not verified.
  • Bears on: Q1 growth rate, Q3 demand, Q8 benchmarks.
  • Unused: not yet cited in the argument. It is the one available quality-adjustment on the throughput gains the rest of this section reports, and its duplication finding is the same phenomenon the log tracks as duplicate bug reports [→ XBOW] and homogenized output [→ idea diversity].
  • Links: report page · PDF
  • Status: vendor, verified as to figures — the composition table, the duplicate-block table, and the quotes were read from the extracted text of the PDF, retrieved 2026-07-26. The vendor sells tooling premised on this being a problem, and no part of the analysis has been independently replicated.

Measurement and benchmark validity

Q8 asks which benchmarks predict real value, and it is the only one of the eight questions where the relevant literature is about the instruments rather than about AI. These entries audit the measurements the rest of the log depends on. They are collected rather than distributed across the domains because the failure modes are general: a score can move without any capability moving, a ranking can invert under a different scoring rule, and a proxy can be optimized while the objective it stands in for does not budge.

Two entries auditing the optimization benchmarks specifically stay in the Algorithms section, because that is what they audit [→ benchmark reliability, PERFOPT]. The FrontierMath funding-disclosure episode is recorded inside that entry for the same reason [→ FrontierMath]. Grace’s selection warning is the oldest statement of the general problem and sits with her survey [→ Grace].

Kapoor and co-authors: AI agents that matter (2024)

Independent (academic, Princeton). The general methodological critique of agent benchmarking, and the reason to distrust a leaderboard position as evidence of anything. It is the parent of the specific audits elsewhere in the log.

  • Accuracy-only evaluation distorts what gets built. “First, there is a narrow focus on accuracy without attention to other metrics. As a result, SOTA agents are needlessly complex and costly, and the community has reached mistaken conclusions about the sources of accuracy gains.”
  • Holdout sets are often absent, so overfitting is invisible. “Third, many agent benchmarks have inadequate holdout sets, and sometimes none at all. This has led to agents that are fragile because they take shortcuts and overfit to the benchmark in various ways.”
  • Two audiences get conflated, which is exactly the Q8 question. “Second, the benchmarking needs of model and downstream developers have been conflated, making it hard to identify which agent would be best suited for a particular application.”
  • Reproducibility. “Finally, there is a lack of standardization in evaluation practices, leading to a pervasive lack of reproducibility.”
  • Why it is load-bearing for this log rather than a caveat. Almost every capability figure recorded here is a benchmark score, and the first point above says such scores conflate capability with cost and scaffold sophistication — which is the same confound Naptime measured at twentyfold [→ Naptime] and PERFOPT measured across agent frameworks [→ PERFOPT]. Three independent routes to one conclusion. This synthesis is the log’s.
  • Dates: arXiv 2024-07-01.
  • Bears on: Q1 growth rate, Q5 returns, Q8 benchmarks.
  • Unused: not yet cited in the argument. It should be the citation behind the argument’s cross-cutting limitations paragraph, which currently asserts these problems without a source.
  • Links: arXiv 2407.01502
  • Status: verified-abstract — all four quotes checked against the arXiv abstract, retrieved 2026-07-26.

The SWE-bench illusion: memorization rather than reasoning (2025)

Independent (academic, Purdue and Microsoft authors). The strongest contamination critique of the benchmark most lab announcements quote, and it supplies its own out-of-distribution control rather than only raising the possibility.

  • The claim, hedged by the authors themselves. “We present empirical evidence that performance gains on SWE-Bench-Verified may be partially driven by memorization rather than genuine problem-solving.” The word is “partially,” and the entry keeps it.
  • The headline diagnostic, with the control that makes it a diagnostic. “We show that state-of-the-art models achieve up to 76% accuracy in identifying buggy file paths using only issue descriptions, without access to repository structure. This performance is merely up to 53% on tasks from repositories not included in SWE-Bench, pointing to possible data contamination or memorization.” Locating a bug without seeing the code is only possible if the answer is already known.
  • A second diagnostic pointing the same way. “Similar patterns are also observed for the function reproduction task, where the verbatim similarity is much higher on SWE-Bench Verified than on other similar coding benchmarks (up to 35% consecutive 5-gram accuracy on SWE-Bench Verified and Full, but only up to 18% for tasks in other benchmarks).”
  • What it does to the rest of the log. SWE-bench scores are the usual public evidence for rapid agentic coding progress, and this says an unknown share of the level is memorization. It does not follow that the trend is spurious, since contamination would have to be growing over time to produce a false slope — but nothing here establishes that it is not. This reading is the log’s.
  • Dates: arXiv 2025-06-14 (v4 2025-12-01).
  • Bears on: Q1 growth rate, Q7 incidence, Q8 benchmarks.
  • Unused: not yet cited in the argument. It belongs beside any use of SWE-bench-derived progress claims [→ SWE-bench].
  • Links: arXiv 2506.12286
  • Status: verified — the abstract of v4 read verbatim on arXiv, retrieved 2026-07-26.

Sakana on kernel benchmarks: a vendor conceding exploitable loopholes (2025)

Vendor (Sakana AI). Included because it is a vendor stating in a primary document that the previous generation of CUDA-kernel speedup claims was measured on a gameable harness. That makes it evidence about the reliability of kernel-speedup figures generally, which several entries here report.

  • The concession about the existing benchmarks. “existing kernel generation benchmarks suffer from exploitable loopholes and insufficient diversity in testing conditions, hindering true generalization assessment.”
  • What they built in response. “we introduce robust-kbench, a new benchmark for rigorous evaluation of kernel performance and correctness across varied scenarios.”
  • The result is stated qualitatively, and the absence of a multiple is the notable part. “Evaluated on robust-kbench, our approach produces CUDA kernels outperforming torch implementations for practical applications, including forward and backward passes.” No headline speedup factor appears, which is a marked change from the earlier generation of claims in this area.
  • Why it matters for reading the kernel results in this log. TTT-Discover’s cross-hardware kernel gains are reported against human leaderboard submissions rather than against a torch baseline [→ TTT-Discover], which is a stronger comparator, but the general lesson stands: a kernel speedup is a measurement on a harness, and harnesses in this area have been shown to be gameable by the people building the agents. This framing is the log’s.
  • Dates: arXiv 2025-09-16; the landing page for the earlier work now resolves to this paper.
  • Bears on: Q1 growth rate, Q2 autonomy, Q8 benchmarks.
  • Unused: not yet cited in the argument. It is the vendor-side corroboration of the argument’s claim that raw multipliers do not travel.
  • Links: arXiv 2509.14279 · project page
  • Status: verified-abstract — quotes checked against the arXiv abstract, retrieved 2026-07-26. The earlier claims this paper supersedes, and the independent replication that reportedly found them overstated, are secondhand and are deliberately not quoted here.

GDPval: a benchmark built to predict economic value, with its authors’ limits (2025)

Vendor (OpenAI). The current best-funded attempt to connect a benchmark score to economic value, and a useful source on its own limitations. It measures professional deliverables rather than research, so its relevance is as a template for what a value-predictive benchmark looks like.

  • Construction and coverage. “GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP. Tasks are constructed from the representative work of industry professionals with an average of 14 years of experience. We find that frontier model performance on GDPval is improving roughly linearly over time, and that the current best frontier models are approaching industry experts in deliverable quality.”
  • Denominators. 1,320 tasks across 44 occupations, with a 220-task open-sourced gold subset across 9 sectors.
  • The headline is a win-or-tie rate on the subset, not a win rate on the whole. “Claude Opus 4.1 was the best performing model on the GDPval gold subset,” with “47.6% of deliverables…graded as better than (wins) or as good as (ties) the human deliverable.” Both qualifications matter and both are routinely dropped.
  • The authors’ own limitations, which are the reason this is a Q8 entry rather than a Q1 one. Tasks “are precisely-specified and one-shot, not interactive”; the evaluation focuses on “self-contained knowledge work”; the dataset is an “initial cut” rather than comprehensive; and it cannot capture “extensive tacit knowledge” or work requiring “communication between individuals.” Research is interactive, ill-specified, and tacit-knowledge-heavy, so a GDPval score is not a proxy for research capability.
  • Dates: arXiv 2025-10-05.
  • Bears on: Q1 growth rate, Q2 autonomy, Q8 benchmarks.
  • Unused: not yet cited in the argument. It is the best available model of what the argument asks for in Q8 — a benchmark carrying a deployment-setting denominator — and simultaneously an illustration of how far even that is from measuring research.
  • Links: arXiv 2510.04374
  • Status: verified — abstract and metadata from the arXiv listing and the task counts, win-or-tie rate, and limitations from the arXiv HTML of v1, retrieved 2026-07-26. Produced by a lab evaluating frontier models including competitors’.

AI-discovered drugs: the proxy clears, the objective does not (2024)

Vendor (Boston Consulting Group authors), peer-reviewed journal article. Clinical-trial outcomes for molecules discovered by AI-native biotech companies. This is the cleanest case in the log of a measurable proxy being optimized while the objective it stands for does not move, which is the Q8 failure mode in its purest form.

  • Both phases, with the authors’ own sample caveat in the same sentence. “In Phase I we find AI-discovered molecules have an 80-90% success rate, substantially higher than historic industry averages. This suggests, we argue, that AI is highly capable of designing or identifying molecules with drug-like properties. In Phase II the success rate is ∼40%, albeit on a limited sample size, comparable to historic industry averages. Our findings highlight early signs of the clinical potential of AI-discovered molecules.”
  • The structure of the result is the finding. Phase I tests safety and tolerability, which track drug-likeness — a property with computable proxies to optimize against. Phase II tests whether the drug works in patients, which has no such proxy. The advantage is large where a proxy exists and absent where it does not. That is the same pattern as cheap-verifier dominance across this log’s three domains, arriving from pharmacology. The reading is the log’s; the authors say only “comparable to historic industry averages.”
  • The denominator is not established here and must not be invented. The abstract says only “a limited sample size.” Secondary sources give molecule counts between about twenty and about seventy; none is verified, so no rate from this entry should be quoted with an N.
  • Dates: published in Drug Discovery Today 29(6), 2024.
  • Bears on: Q1 growth rate, Q2 autonomy, Q5 returns, Q8 benchmarks.
  • Unused: not yet cited in the argument. It is the strongest external-domain illustration of the argument’s verification claim, and it carries the external-validity discount that biology is not cyber, math, or algorithms.
  • Links: DOI 10.1016/j.drudis.2024.104009
  • Status: verified-abstract — the abstract was retrieved through the Europe PMC record for the DOI, retrieved 2026-07-26, because the publisher blocks automated fetching; the body has not been read and no denominator has been confirmed. The authors are consultants with commercial interests in AI-native biotech.

Stack Overflow developer survey: adoption high, trust low (2025)

Independent (Stack Overflow), self-selected survey. The practitioner-reported version of the Q8 question: why a capable-looking tool can fail to produce value. Weak evidence by the log’s own ranking, included because the “almost right” figure names a specific mechanism that the measured studies corroborate.

  • Adoption. “84% of respondents are using or planning to use AI tools in their development process,” and “51% of professional developers use AI tools daily.”
  • Trust, with both sides. “46% of developers actively distrust the accuracy of AI tools” against 33% who trust it and 3% who are “highly trusting.”
  • The mechanism, and it is the one the METR trial measured. 66% encounter “AI solutions that are almost right, but not quite,” and 45.2% report that “Debugging AI-generated code is more time-consuming.” METR’s randomized trial found exactly this overhead — reviewing and correcting output offsetting the time saved generating it [→ METR RCT] — so the survey’s mechanism has an experimental counterpart even though the survey itself cannot establish it.
  • Denominator and its limits. 33,662 responses on the overall usage question. Self-selected sampling, so these are not population estimates and the log’s evidence ranking puts stated perceptions at the bottom.
  • Dates: published 2025; field dates are not stated on the page read.
  • Bears on: Q1 growth rate, Q3 demand, Q8 benchmarks.
  • Unused: not yet cited in the argument. Its use is to name the mechanism behind the METR result in practitioners’ own terms, not to establish anything.
  • Links: survey AI section
  • Status: unverified as measurement, verified as to figures — the percentages and the response count were read from the survey page, retrieved 2026-07-26. A self-selected online survey of stated perceptions, with no counterfactual.

Simple baselines match code evolution, so the machinery may not be what works (2026)

Independent (academic, Gideoni, Risi, and Gal). An attribution audit of the evolutionary-coding-agent literature, including a direct replication of nine AlphaEvolve problems. It is the only entry in this log that asks whether the search method credited with a mathematical result is what produced it, which is the attribution question every AI-discovery claim in the Math and Algorithms sections rests on.

  • The claim, across three domains. The paper tests “two simple baselines over three domains: finding better mathematical bounds, designing agentic scaffolds, and machine learning competitions,” and reports that the simple baselines “match or exceed much more sophisticated methods in all three domains.”
  • The replication, with its denominator. Nine problems from the AlphaEvolve paper were used as case studies [→ AlphaEvolve mathematics]. Randomly sampling programs from an LLM matched AlphaEvolve on two problems, and matched or improved on the strong open-source baseline ShinkaEvolve on eight of the nine.
  • The mechanism claim is the important part, and it relocates the credit. For mathematical bounds, the search space and the domain knowledge placed in the prompt are what chiefly determine performance, with the evolutionary pipeline secondary; the authors conclude the primary challenge is designing good search spaces rather than the search itself. If that is right, an AI-improved bound is substantially a human-specified-search-space result, which is the same “expertise moved to the harness” pattern the algorithms domain shows [→ TTT-Discover, PERFOPT].
  • What it does not claim. It does not say the bounds were not improved, nor that LLMs contributed nothing — random sampling from an LLM is still using the model. It says the sophistication of the scaffold is not where the gain comes from. Conflating the two would overstate it, and this distinction is the log’s.
  • Scope caveat. Nine problems out of 67, and the paper is a preprint whose peer-review status is unclear: closely-titled versions appear on OpenReview as both “Simple Baselines are Competitive with Code Evolution” and “Random Baselines for Simple Code Problems are Competitive with Code Evolution,” the latter listed against NeurIPS 2025. This log has not established which is the version of record, and the two titles differ in how strong a claim they make.
  • Dates: arXiv 2026-02-18; the OpenReview and NeurIPS 2025 listings of the closely-titled versions are not separately dated here.
  • Bears on: Q1 growth rate, Q4 expertise, Q8 benchmarks.
  • Unused: not yet cited in the argument. It belongs wherever the argument credits an AI system with a mathematical or algorithmic improvement, as the discount on attributing that improvement to the search method.
  • Links: arXiv 2602.16805 · OpenReview · closely-titled NeurIPS version
  • Status: unverified — the three-domain claim, the nine-problem replication counts, the search-space conclusion, and the 2026-02-18 date are taken from the arXiv listing and search-surfaced summaries, retrieved 2026-07-26. Neither the paper body nor the OpenReview versions have been read, and no figure is quoted. Verify before the argument leans on it.

Expertise and demand

Almost nothing in this section is about cyber, math, or algorithms. It is here because Q4 lost its headline finding to a fabrication [→ Toner-Rodgers] and Q3 had no occupation-level evidence at all, so the argument was reasoning about expertise and demand from domain anecdote. This literature measures the interaction between AI and user ability directly, often with randomization, and it measures labor-market outcomes with a counterfactual — on tasks and occupations that are mostly not research.

Every entry carries an external-validity discount and should be cited with it: a consultant writing a memo is not a mathematician attacking an open problem, a freelance copywriter is not a security researcher, and the reachable zone in one setting says little about the other. Two entries carry a smaller discount than the rest, because software developers are among the occupations they cover [→ canaries, and the two Copilot studies in the algorithms section]. The section was previously titled “outside the three domains,” which stopped being accurate once those were added.

The section’s most useful property is that the studies disagree about the sign, and the disagreement is structured rather than noisy: it tracks whether the task has a checkable answer and how good the starting point was. The synthesis below collects the designs in one place [→ experimental evidence].

Dell’Acqua and co-authors: the jagged technological frontier (2023)

Independent (academic, preregistered field experiment), peer-reviewed. 758 Boston Consulting Group consultants randomized to no AI, GPT-4, or GPT-4 with a prompt-engineering overview, after a baseline performance measurement. The source of the “jagged frontier” term the cyber entries use.

Grouped bar chart: bottom-half skilled participants rise from 4.37 to 5.72, a 31 percent gain, and top-half participants from 5.34 to 5.82, an 11 percent gain.

Task scores for bottom-half and top-half skilled participants, with and without AI, on inside-the-frontier tasks.

The paper’s Figure 4, for the inside-the-frontier tasks. Bottom-half skilled participants score 4.37 at baseline and 5.72 with AI, a 31% gain; top-half participants score 5.34 and 5.82, an 11% gain. Both groups improve and the gap between them narrows. The paper’s outside-the-frontier experiment, reported separately, runs in the other direction.

  • Inside the frontier the effect is large and positive. Across 18 realistic consulting tasks, subjects with AI completed 12.2% more tasks, 25.1% more quickly, at significantly higher quality.
  • Outside it the effect reverses sharply. On a complex managerial task chosen to sit beyond AI’s capability, subjects using AI were 19% less likely to reach a correct solution than those without it. Same workers, same workflow, similar apparent difficulty.
  • This is the cleanest experimental statement of a reach boundary. A task-level discontinuity in the sign of the effect, established by randomization, is what apple-picking and task replacement both predict and what uniform acceleration does not. It does not discriminate between the two, because both posit a boundary; they differ on what happens at it.
  • Gains were largest for the lowest performers. The below-average consultants gained most, so within the frontier the effect was skill-compressing.
  • Idea diversity fell. The paper documents a compression in the diversity of ideas among AI-using consultants, which is the same homogenization mechanism the recombinant-innovation model predicts [→ recombinant innovation].
  • Dates: field experiment conducted 2023 with GPT-4; HBS working paper September 2023; published in Organization Science 2026-03-01.
  • Bears on: Q4 expertise, Q5 returns, Q7 incidence.
  • Unused: not yet cited in the argument, though the argument already uses the word “jagged”. If the term is used, this is where it comes from and it should be cited.
  • Links: Organization Science · HBS PDF
  • Status: verified-abstract — headline figures checked against the published abstract and the HBS working-paper PDF, retrieved 2026-07-26. Figure reproduced from the source.

Brynjolfsson, Li, and Raymond: generative AI at work (2023)

Independent (academic, staggered field rollout). About 5,000 customer-support agents at a software firm, using the phased deployment of an AI assistant for identification. The most-cited evidence that AI compresses the within-occupation skill distribution.

Two panels of point estimates with confidence intervals: gains fall monotonically from about +0.27 log points in the lowest skill quintile to about zero in the highest, and from about +0.38 for zero-tenure workers to about zero above twelve months.

Change in log resolutions per hour after AI deployment, by pre-AI worker skill quintile (panel A) and by tenure (panel B).

The paper’s Figure 5. Panel A groups agents into quintiles of pre-AI skill and panel B by tenure at deployment; both plot the change in log resolutions per hour following deployment. The gradient is monotone in both panels, and in each the top group’s estimate is indistinguishable from zero.

  • Average productivity rose 14%, and the gain was concentrated among the least skilled. In the authors’ words, access to the tool “increases productivity, as measured by issues resolved per hour, by 14% on average, including a 34% improvement for novice and low-skilled workers but with minimal impact on experienced and highly skilled workers.” The sample is 5,179 agents.
  • “Minimal impact” is visible as zero in the figure. The top skill quintile and workers with more than twelve months’ tenure both have point estimates indistinguishable from zero.
  • The implied mechanism is knowledge transfer. The assistant surfaced the practices of high performers to everyone else, which raises the floor without moving the ceiling — a floor-raise, not a frontier-shift.
  • Why the sign matters here. If AI mainly substitutes for expertise the researcher does not have, it should widen participation in discovery without accelerating the frontier. That is closer to apple-picking’s reachable zone than to human replacement. But the setting is a routine task with a known-good answer, which is the least research-like setting imaginable.
  • Dates: NBER working paper 31161 issued 2023-04-20; deployment data from 2020–2021; subsequently published in the Quarterly Journal of Economics.
  • Bears on: Q3 demand, Q4 expertise.
  • Unused: not yet cited in the argument.
  • Links: NBER w31161
  • Status: verified — NBER working paper 31161 (April 2023, revised November 2023) read directly, retrieved 2026-07-26; the 14%, 34%, and 5,179 figures are quoted from its abstract. Figure 5 reproduced from the source. This supersedes an earlier unverified status that flagged the 14% and 34% as unchecked.

Otis and co-authors: the uneven impact on Kenyan entrepreneurs (2024)

Independent (academic, randomized field experiment). 640 Kenyan small-business owners randomized to a GPT-4-powered business adviser over WhatsApp, tracked for five months. Included because it finds the opposite heterogeneity to the customer-support study, on a more open-ended task.

Four point estimates with confidence intervals in standard deviations: full sample +0.04, initial low performers minus 0.08, initial high performers +0.16, and the heterogeneous effect +0.23.

Effect of the AI assistant on a standardized business-performance index: full sample, initial low performers, initial high performers, and the difference.

The paper’s Figure 3. The outcome is a standardized business-performance index. Panel A is the average treatment effect, 0.04 s.d. with a confidence interval spanning zero; panels B and C split the sample by initial performance, at −0.08 s.d. for low performers and +0.16 for high; panel D is the difference between them, 0.23 s.d. Estimates control for pre-treatment performance, time and stratum fixed effects, and covariates selected by double-LASSO.

  • No average effect, and a large gap by baseline ability. The authors cannot reject a null average treatment effect on revenues and profits, but the effect for baseline low performers is nearly 0.25 standard deviations below that for high performers.
  • High performers gained more than 15%; low performers lost nearly 10%. AI access actively harmed the weaker half.
  • The mechanism is selection among suggestions, not different suggestions. The paper finds the divergence does not come from differences in questions asked or advice received, but from which advice entrepreneurs chose to implement. Judgement about which output to trust is the scarce complement.
  • This is the single most important contrast in this section. Where the task has a checkable right answer, AI compresses the skill distribution; where it is open-ended and the user must filter, AI widens it. Research is the open-ended case. That points toward AI raising the return to expert judgement in exactly the settings this project cares about — and it is also the mechanism the METR developer RCT describes, where accepting AI output uncritically was the cost [→ METR RCT].
  • Dates: five-month RCT; HBS working paper 24-042, 2024; pre-published online in Management Science 2026-07-10; a practitioner summary appeared in MIT Sloan Management Review, Summer 2026.
  • Bears on: Q4 expertise, Q7 incidence.
  • Unused: not yet cited in the argument. With the customer-support study it forms the pair that makes the Q4 row defensible again after the Toner-Rodgers withdrawal.
  • Links: HBS working paper PDF · HBS listing · Berkeley Haas summary
  • Status: verified — full working-paper PDF read, retrieved 2026-07-26. The null average effect, the ±10%/15% subsample figures, and the selection mechanism are checked against the body; the panel values in the figure are the paper’s own printed coefficients. Figure 3 reproduced from the source. Note that the gap is 0.23 s.d. as printed in Figure 3, which the text rounds to “nearly 0.25.”

Does generative AI narrow education-based productivity gaps? (2026)

Independent (academic, randomized online experiment). 1,174 adults aged 25–45 with heterogeneous education, randomized to complete an incentivized business problem-solving task with or without a generative-AI assistant. Designed specifically to estimate the between-education-group gap, which within-firm studies cannot.

  • AI closed about three quarters of the education-based productivity gap. Higher-education participants outperformed lower-education participants by 0.548 standard deviations without AI; with AI the gap fell to 0.139.
  • Everyone gained, but the low-education group gained much more. So this is compression, agreeing with the customer-support study and disagreeing with the Kenya experiment.
  • The compression is in task execution, not in capability. Education gaps persisted in a follow-up exercise without AI, which the authors read as AI relaxing cognitive constraints rather than transferring human capital. The distinction matters for Q4: borrowed capability disappears when the tool does.
  • Design note. Conducted outside firms deliberately, because organizational selection compresses educational heterogeneity and makes within-firm estimates unrepresentative of the population gap. That is a real advantage over the other entries here, offset by the task being an artificial exercise.
  • Dates: NBER working paper 34851, issued February 2026 (2026-02-13).
  • Bears on: Q4 expertise.
  • Unused: not yet cited in the argument.
  • Links: NBER w34851
  • Status: verified-abstract — all figures checked against the paper’s abstract, retrieved 2026-07-26.

Doshi and Hauser: AI raises individual creativity and lowers collective diversity (2024)

Independent (academic, peer-reviewed in Science Advances). Writers randomized to generative-AI story ideas. Included because it measures the aggregate-level cost of a tool that helps each user individually, which is the shape of the duplication problem in the discovery domains.

  • Individually more creative, collectively more alike. Stories written with AI assistance were rated more creative on average, while the diversity of the resulting body of work fell by roughly 10%.
  • This is the streetlight and stepping-on-toes mechanism, measured. The recombinant-innovation model predicts that shared systems concentrate search in the same regions and make researchers converge on the same suggestions [→ recombinant innovation]; XBOW’s duplicate rate is the same phenomenon in cyber [→ XBOW]. This is the cleanest experimental demonstration of it, in a domain where output diversity can be measured directly.
  • Why it bears on apple-picking. Every searcher using the same model reaches for the same apples. Under apple-picking that produces duplicates and a depleting reachable zone; under task replacement it does not obviously produce anything. The prediction is shared with the recombinant model, so the evidence supports the family rather than the specific theory.
  • Scope. Creative writing, not research, and diversity of stories is not diversity of ideas in a technical field.
  • Dates: published in Science Advances, July 2024. The exact publication date has not been confirmed here.
  • Bears on: Q5 returns, Q7 incidence.
  • Unused: not yet cited in the argument.
  • Links: Science Advances
  • Status: unverified — the primary page is behind a bot check and could not be fetched, so the figures and the date come from secondary summaries, retrieved 2026-07-26. Do not cite until the primary is read.

Noy and Zhang: ChatGPT compresses the writing productivity distribution (2023)

Independent (academic, preregistered randomized experiment), peer-reviewed in Science. The cleanest randomized evidence that generative AI helps lower-ability workers more (Noy and Zhang 2023). It is also a case where the working paper and the published version report different figures, which the entry records rather than resolving.

  • The working-paper version, in standard deviations. “In a preregistered online experiment, we assign occupation-specific, incentivized writing tasks to 444 college-educated professionals, and randomly expose half of them to ChatGPT. Our results show that ChatGPT substantially raises average productivity: time taken decreases by 0.8 SDs and output quality rises by 0.4 SDs.”
  • The compression result and its mechanism, which is the Q4-relevant sentence. “Inequality between workers decreases, as ChatGPT compresses the productivity distribution by benefiting low-ability workers more. ChatGPT mostly substitutes for worker effort rather than complementing worker skills, and restructures tasks towards idea-generation and editing and away from rough-drafting.”
  • The published version, in percentages and with a different sample count. “we assigned occupation-specific, incentivized writing tasks to 453 college-educated professionals and randomly exposed half of them to ChatGPT. Our results show that ChatGPT substantially raised productivity: The average time taken decreased by 40% and output quality rose by 18%.”
  • The version discrepancy, recorded per this log’s rule. The March 2023 working paper reports 444 participants and 0.8 / 0.4 standard deviations; the July 2023 Science version reports 453 participants and 40% / 18%, and drops the “substitutes for worker effort” sentence from the abstract in favor of one on persistence of adoption. Cite whichever version you actually mean.
  • Substitution rather than complementarity is the finding to carry forward. If AI substitutes for effort rather than complementing skill, then it should compress outcomes wherever the task has a checkable answer and do nothing for the skill itself — which is what the education-gap experiment found when it tested for retention [→ education gap]. Writing a memo is a long way from attacking an open problem, and the external-validity discount for this whole section applies with full force.
  • Dates: working paper 2023-03-02, preregistered at the AEA RCT Registry; published in Science 381, 2023-07-14.
  • Bears on: Q1 growth rate, Q4 expertise, Q8 benchmarks.
  • Unused: not yet cited in the argument. With the customer-support study and the education-gap experiment it is the third randomized compression result, which is what makes the Kenyan divergence finding informative rather than anomalous.
  • Links: working paper PDF · Science
  • Status: verified for the working paper, whose quotes were read from the PDF, retrieved 2026-07-26. The published figures were read from an aggregator’s record of the DOI rather than from the publisher page, which blocks automated fetching, so they are verified-abstract at one remove and should be re-checked against Science before being leaned on.

Hui, Reshef, and Zhou: freelance demand fell, and top freelancers fell hardest (2023)

Independent (academic), difference-in-differences. Employment and earnings for freelancers on a large online platform around the release of ChatGPT. It is the closest thing in the log to a measured demand effect on knowledge work, and its heterogeneity result points against the complementarity story the log’s other expertise entries suggest.

  • The headline estimates, with standard errors. “Following the release, the monthly number of jobs on the platform for freelancers in more affected occupations decreases by 2% (s.e.=0.004), and total monthly compensation decreases by 5.2% (s.e.=0.016).” On the extensive margin, freelancers are “1.2 percentage points” less likely to receive any job in a given month, “which is approximately a 10% drop compared” to baseline.
  • The quality gradient runs the wrong way for complementarity, and the authors are explicit. “Exploring the heterogeneity by freelancers’ employment history, we do not find evidence that high-quality service, measured by their past performance and employment, moderates the adverse effects on employment. In fact, we find suggestive evidence that top freelancers are disproportionately affected by AI. These results suggest that in the short term generative AI reduces overall demand for knowledge workers of all types, and may have the potential to narrow gaps among workers.”
  • Sample and window. “We restrict our attention to the period from January 2022 through April 2023… Our final data set consists of 92,547 freelancers.” For scale: “on average, a freelancer starts a job once every three months, for an average monthly pay of $171.”
  • The authors’ own horizon caveat, which should travel with the figures. “Notably, in this paper we provide novel, preliminary evidence on the short-term effects of generative AI. However, the long-term implications may be significantly different, and it is unclear how our findings extend to longer time horizons.”
  • A disclosed data limitation that biases the estimate toward zero. “our sample only includes freelancers with active profiles at the time of obtaining the data, as we do not observe terminated accounts,” and they observe only “a snapshot of the freelancer pool as it was in April 2023.” Freelancers who left entirely are missing, so the measured decline is a lower bound. The direction of the bias is the log’s inference from their statement.
  • Dates: CESifo working paper 10601, 2023; data window January 2022 to April 2023; published in Organization Science 35(6), 2024.
  • Bears on: Q3 demand, Q4 expertise, Q7 incidence.
  • Unused: not yet cited in the argument. It is the demand-side counterpart the Q3 row asks for, on knowledge work rather than on research, and its top-freelancer result is the sharpest available challenge to the argument’s reading that AI raises the return to expert judgement.
  • Links: CESifo PDF · RePEc listing
  • Status: verified — all quotes read from the CESifo working-paper PDF, retrieved 2026-07-26. The published Organization Science version was not read and its figures may differ; occupational exposure is measured by classification rather than by observed AI use.

Demirci, Hannane, and Zhu: postings for automation-prone freelance work fell 21% (2024)

Independent (academic), difference-in-differences. The companion demand-side study measured on job posts rather than on freelancer outcomes, separating text generation from image generation. Its finding about what survives is the apple-picking-shaped part.

  • Both headline figures, with the comparison group and the window. “Our findings indicate a 21% decrease in the number of job posts for automation-prone jobs related to writing and coding, compared to jobs requiring manual-intensive skills, within eight months after the introduction of ChatGPT… We also find that the introduction of Image-generating AI technologies led to a 17% decrease in the number of job posts related to image creation.”
  • What remains gets harder and better paid. “We show that the reduction in the number of job posts increases competition among freelancers while the remaining automation-prone jobs are of greater complexity and offer higher pay.” A residual that is more complex after the easy work is absorbed is the composition change apple-picking predicts, observed in a labor market rather than in a research domain.
  • One channel is correlational and they say so. “We use Google Trends to show that the more pronounced decline in the demand for freelancers within automation-prone jobs correlates with their higher public awareness of ChatGPT’s substitutability.”
  • Dates: CESifo working paper 11276, 2024; the window is the eight months after the ChatGPT release; accepted at Management Science.
  • Bears on: Q3 demand, Q5 returns, Q7 incidence.
  • Unused: not yet cited in the argument. Together with the freelancer-outcome study it gives Q3 two quasi-experimental estimates where the log currently records that it has none.
  • Links: RePEc listing
  • Status: verified-abstract — the quotes were checked against the CESifo record, retrieved 2026-07-26. No sample size appears in the abstract and the working-paper PDF was not opened, so do not attach a denominator to this entry.

Brynjolfsson, Chandar, and Chen: entry-level employment in AI-exposed occupations (2025)

Independent (Stanford Digital Economy Lab), administrative payroll data. The most-cited labor-market evidence on entry-level displacement, and the entry in this section with the least external-validity discount, because software developers are among the exposed occupations (Brynjolfsson, Chandar, and Chen 2025). It still measures a labor market rather than a research frontier.

  • The headline, with its control and the null for experienced workers. “Using high-frequency administrative data from ADP, we document six facts characterizing labor market shifts following the widespread adoption of generative AI. Early-career workers (ages 22-25) in AI-exposed occupations experienced 16% relative employment declines, controlling for firm-level shocks, while employment for experienced workers remained stable. Adjustments occur primarily via employment rather than compensation, with employment changes concentrated in occupations where AI automates rather than augments labor. Results are robust to excluding technology firms and occupations that are remotable.”
  • The automation-versus-augmentation split, which is the finding that discriminates. “Fact 3: Entry-level employment has declined in applications of AI that automate work, with muted changes for augmentation.” Their exposure measure comes from a vendor: “we use data on generative AI usage from the Anthropic Economic Index,” which “provides an estimate of the share of queries that pertain to each occupation” and classifies queries as “automative,” “augmentative,” or neither. That dependency is worth naming, since the key split rests on a vendor’s classification of its own traffic.
  • The margin of adjustment. “Fact 5: Labor market adjustments are visible in employment more than compensation… The findings indicate less divergence in compensation compared to employment across more and less exposed occupations.”
  • Breadth across the exposure distribution, read from their appendix figure rather than quoted. Close to 70% of occupations in the least-exposed quintile see rising early-career employment over October 2022 to September 2025, against under half in the most-exposed quintile.
  • The authors’ own epistemic hedge, and their choice of the word “facts”. “These six facts provide early large-scale evidence consistent with generative AI disproportionately impacting entry-level workers in the American labor market.” The abstract says “consistent with,” not “caused by,” and the paper calls them facts rather than estimates.
  • No exact sample size exists to cite. The paper reports only “monthly, individual-level payroll records through September 2025, encompassing millions of workers across tens of thousands of firms.” Do not invent a denominator for it.
  • An earlier draft reported a different figure. Versions circulating from August 2025 give 13% where the November version gives 16%, so pin the version when quoting.
  • How it bears on this project. The log’s Q3 gap asks for occupation-level causal evidence on security researchers, mathematicians, or ML engineers. This is not that, but it is the nearest available: an exposed-occupation set that includes software development, an age gradient, and an automation-versus-augmentation split that maps onto the task-replacement prediction. The cyber workforce survey reports the same age pattern from self-report [→ cyber labour], which is weak corroboration from an independent direction.
  • Dates: version dated 2025-11-13; data window October 2022 to September 2025.
  • Bears on: Q3 demand, Q4 expertise, Q7 incidence.
  • Unused: not yet cited in the argument. It is the strongest entry available for the Q3 row and should be used there with the exposure-measure dependency stated.
  • Links: working paper PDF · landing page
  • Status: verified — all quotes read from the working-paper PDF, retrieved 2026-07-26. Not peer-reviewed; exposure is measured from a vendor-supplied classification; and no exact worker count is published.

Aghion and co-authors: French firms that adopted AI grew (2025)

Independent (academic, peer-reviewed proceedings), difference-in-differences. Firm-level AI adoption in France. It points the opposite way from the freelancer and payroll studies above, and the reason is worth holding onto: it measures a different era of AI and a different unit.

  • The whole abstract, since each of the four findings carries. “Using French firm-level data on AI adoption from 2017–2020, we find that, first, firms adopting AI are larger and more productive and skill intensive. Second, difference-in-difference estimates reveal an increase in firm-level employment and sales after AI adoption, suggesting that the induced productivity gains allow firms to grow and outweigh potential displacement effects. Third, occupations classified in recent work as substitutable with AI expand. Fourth, AI usage is a relevant dimension of heterogeneity in the labor demand response: We find positive employment growth for certain uses (e.g., information and communications technology security) and negative for others (e.g., administrative processes).”
  • The scope limit is decisive and is easy to miss. The adoption window ends in 2020, so “AI” here is machine learning and analytics rather than generative models or agents. The first finding is explicitly a selection statement rather than an effect. Setting this against the 2023–2025 studies above is therefore not a contradiction to be resolved but two different technologies measured in two different periods. This point is the log’s.
  • The fourth finding is the transferable one. Employment rose for security uses and fell for administrative ones, within the same firms and period. Incidence by use case rather than by occupation or firm is the cut the argument’s Q7 row wants, and this is the only entry in the log that makes it on employment data.
  • Dates: adoption data 2017–2020; published in AEA Papers and Proceedings 115, May 2025.
  • Bears on: Q3 demand, Q5 returns, Q7 incidence.
  • Unused: not yet cited in the argument. Its use is as the pre-generative baseline for the demand question, and as the one measured instance of incidence varying by application rather than by worker.
  • Links: AEA article page
  • Status: verified-abstract — the abstract checked against the AEA article page, retrieved 2026-07-26. A four-page proceedings note; the body and identification checks were not read.

Zhao and co-authors: AlphaFold barely changed who collaborates (2025)

Independent (academic). An adopter-versus-non-adopter study of structural biologists around AlphaFold, testing the common claim that AI bridges disciplines. It is here as a negative result on the composition channel, in the one domain where an AI tool has most plainly changed practice.

  • The null, with its denominators. “By analyzing 1,247 AlphaFold-related papers and 7,700 authors from Scopus, we employ bibliometric analysis and causal inference to compare interdisciplinary collaboration between AlphaFold adopters and non-adopters. Contrary to the widespread belief that AI facilitates interdisciplinary collaboration, our findings show that AlphaFold increased structural biology-computer science collaborations by just 0.48%, with no measurable effect on other disciplines.”
  • The mechanism they propose, which is a substitution story. “AI creates interdisciplinary collaboration demands with specific disciplines due to its technical characteristics, but this demand is weakened by technological democratization and other factors. These findings demonstrate that artificial intelligence (AI) alone has limited efficacy in bridging disciplinary divides or fostering meaningful interdisciplinary collaboration.”
  • Why a null is worth an entry. If a widely-adopted AI tool democratizes a capability, the researcher who would have supplied that capability is no longer needed as a collaborator — so a null on collaboration is consistent with a large effect on practice. That reading is the authors’ mechanism, and it is the same shifting-rather-than-rising pattern the argument’s Q4 row reports across all three domains.
  • Dates: arXiv 2025-08-18 (v2 2025-10-27).
  • Bears on: Q3 demand, Q4 expertise, Q7 incidence.
  • Unused: not yet cited in the argument. It is a rare measured null on a composition question, in a domain where AI’s contribution is not in dispute.
  • Links: arXiv 2508.13234
  • Status: verified — abstract, authors, and dates checked against the arXiv listing, retrieved 2026-07-26. The abstract asserts causal inference without naming a design, and the body was not read, so treat the identification as quasi-experimental at best.

Syntheses assembled here

Five cross-cutting comparisons that no single source supplies, assembled from the entries above so that the argument and a later reader can cite them once instead of rebuilding them. Every entry in this section carries the derived status: nothing in it is quoted as though a source had said it, every input is named by anchor, and each is reproducible from those inputs. A sixth derived entry, the comparison of efficiency rates across domains, sits with the aggregate measures it is built from [→ efficiency rates].

Four of them exist because the same defects recur across a hundred-odd entries and are invisible one entry at a time. A cost figure means nothing without knowing what the comparable figures are; an autonomy claim means nothing without knowing what the word covered in the other cases; a rate means nothing without a denominator, and most of the rates here do not have one. The fifth is different in kind: it builds a series that did not previously exist, by extracting dated bounds from the exponent database, and it supplies the math domain’s missing pre-AI baseline [→ ANTEDB rates].

How fast analytic number theory’s exponents actually improve (2026)

This log’s synthesis. Not a source: a dated series extracted from the exponent database [→ ANTEDB] and put in the same units as the efficiency curves this project is built on. The math section of the argument asserts that tightened bounds are the right outcome variable for this domain because they give a monotone dated series rather than a binary solved-or-not; this is that series, actually built. As far as this log can establish, nobody had plotted the database’s contents against time before — the database’s own figures plot exponent-pair regions in parameter space, not history.

Thirty small panels in six rows of five. Rows one and two are mu at ten values of sigma, each a nearly flat descending staircase from 1920 to the present, ending between 0.87 and 0.58 of its starting value. Rows three and four are the zero-density exponent A at ten values of sigma: the low-sigma panels have two records each and last moved in 1940, while the high-sigma panels descend in several steps to as little as 0.20 of their starting value, crossing a dashed line marking the density hypothesis. Rows five and six are beta at ten values of alpha, beginning only around 1989, with large drops at small alpha and three panels that never move at all.

Best known value of thirty analytic-number-theory exponents against time, one panel per slice, extracted from ANTEDB.

Each panel is one slice of one exponent: the best value derivable from the literature available in that year, computed by the database’s own solver, with the year taken from the reference that established each input. Every panel runs from 0 to a little above its own earliest value, so the fraction of the panel’s height the line descends is the fraction of the bound that has been removed, and the panels can be compared by eye. Markers are the years the record changed. The dashed line on the \(A\) panels is the density hypothesis \(A\leq2\). Each panel states its first and last value, the ratio between them, the year of the last change, and the number of record changes. The ten slices per family are chosen to include the ones carrying standard names: \(\mu(1/2)\) is the Lindelöf exponent, and \(A(3/4)\) is the slice Ingham bounded in 1940 and Guth–Maynard improved in 2024. Values are as the database records them; the ratios and the choice of slices are this log’s.

Scatter and line plot with the parameter on the horizontal axis from 0 to 1 and the ratio of latest to earliest bound on the vertical axis from about 0.12 to 1.07. A dotted line at 1.0 marks no improvement. Beta rises from 0.25 at small alpha to 1.0 around alpha 0.3 to 0.45. Mu sits between 0.87 and 0.95 across most of sigma before falling to about 0.58 near sigma 1. A declines steadily from 0.97 to 0.20 as sigma approaches 1.

Improvement over the whole record, plotted against the parameter, for all three families.

The same data reduced to one number per slice: the latest bound divided by the earliest, across the full grid of 20 points for \(\mu\), 19 for \(A\) and 19 for \(\beta\). Lower means more of the bound was removed. The dotted line at 1.0 is no improvement; the three \(\beta\) points sitting on it never changed once in the recorded window. The horizontal axis is \(\sigma\) for \(\mu\) and \(A\) and \(\alpha\) for \(\beta\), which are different parameters plotted on one axis for compactness. The grids and the ratios are this log’s arithmetic over the database’s values.

  • How it was built. tools/antedb_extract.py runs against a checkout of the database, calls list_hypotheses(year=Y) to restrict the literature to results published up to each year, asks the database’s own solver for the best bound derivable from that restriction, and walks each derived bound’s dependency tree to attribute it. It writes two files, both vendored: antedb-bounds.csv for the six named slices and antedb-sweep.csv for the parameter grids. tools/sources_figures.py plots from those. Nothing here is quoted from a source, because no source states these series.
  • A subtlety in what the dates mean. The value at year \(Y\) is what the database says was derivable in year \(Y\), not what somebody had written down. Those differ, and the difference is the database’s whole point [→ ANTEDB]: collating relations yields bounds nobody had stated. So this is a curve of available knowledge, not of published claims, and it is the more favorable of the two to plot.
  • The Lindelöf exponent moved 13% in a century. \(\mu(1/2)\) falls from \(5/28 \approx 0.1786\) (van der Corput, 1920) to \(13/84 \approx 0.1548\) (Bourgain, 2017), through fifteen record changes. That is a factor of 0.867 over 97 years, an implied halving time of about 500 years. The conjectured value is 0, so a century of work by many of the strongest analytic number theorists closed about an eighth of the distance.
  • The other five series run from 82 to 1,204 years per halving. \(A(9/10)\) is the fastest at about 82 years, falling 3.6 to 1.5 over 1921–1980 and then flat for 44 years. \(A(3/4)\) is 238 years, \(\mu(3/5)\) 541, \(\mu(7/8)\) 986, \(\mu(3/4)\) 1,204. Every one of these is a halving time measured in centuries.
  • Set against the rates this project’s other entries measure, the gap is three orders of magnitude. Language-model pretraining efficiency halves about every 8 months and ImageNet about every 9 [→ Epoch on LMs, ImageNet]; the fastest physical cost curve in the OWID set, DNA sequencing, halves about every 8.6 months [→ efficiency rates]; ML hardware price-performance takes about 25 months [→ hardware price-performance]. The fastest exponent here is about 40 times slower than the slowest of those, and the slowest is about 1,800 times slower than pretraining efficiency. This comparison is the log’s, and the caveat below limits it.
  • Progress is lumpy, and the flat stretches are long. \(A(3/4)\) has three records in 103 years: Carlson 1921, Ingham 1940, then nothing until Guth–Maynard 2024. Secondary accounts of that gap agree — after Ingham’s 1940 bound, “over the next eighty years, the only improvement to this bound has been small refinements to the o(1) error.” An 84-year plateau in a heavily-studied quantity is the same step-function shape the algorithms domain shows [→ Sherry and Thompson, SAT Museum, Bixby], now in mathematics.
  • A discrepancy found by checking one value against its source paper. The database records Guth–Maynard as \(A(\sigma) \leq 15/(3+5\sigma)\), giving \(A(3/4)=20/9\approx2.222\), while the paper’s headline zero-density consequence is \(N(\sigma,T)\leq T^{30(1-\sigma)/13+o(1)}\), i.e. \(A=30/13\approx2.308\). The database’s recorded form is the stronger of the two. This log has not resolved which the paper’s own sharpest statement is, and the point is recorded because it illustrates what kind of artifact the database is: a live derived object whose entries can be sharper than the abstracts they come from, not a transcription.
  • These six are slices, not the extent of what moves — a correction to an earlier version of this entry. The first version of this entry called six exponents “a small and non-random sample, chosen because the database happens to record them,” which understated the database badly. \(\mu\), \(A\) and \(\beta\) are functions of a parameter, so any point on them is a legitimate dated series, and a sweep across grids of 20, 19 and 19 points finds that nearly all of them move: every \(\mu(\sigma)\) point has between 8 and 18 record changes, every \(A(\sigma)\) point between 2 and 8, and 16 of 19 \(\beta(\alpha)\) points between 2 and 5. The six plotted above were chosen because they carry standard names and standard conjectures, not because they are the only ones with a history.
  • The database is much larger than these three families. It holds 475 dated literature entries across ten hypothesis types, computed here from the entry set: zero-density estimates (155), upper bounds on \(\beta\) (141), exponent pairs (61), upper bounds on \(\mu\) (55), large value estimates (42), large value energy regions (13), plus four smaller types. Two of those types have no history at all in the recorded window — the zero-density energy estimates are three entries all from 1979, and the zeta large value estimate is a single 1978 entry — and the exponent pairs and large-value regions are points and polytopes rather than scalars, so turning them into a series needs a scalarization choice this log has not made. Derived quantities such as the prime-gap exponents, which the database computes from zero-density and energy estimates, are a further untouched family.
  • Progress is strikingly non-uniform along each function, which is a Q7 observation with no AI in it. The second figure is the finding: the century’s improvement ratio varies from 1.00 to 0.20 depending only on where you look. \(A(\sigma)\) improves steadily more as \(\sigma\) approaches 1, ending at 0.20 of its 1921 value at \(\sigma=39/40\) against 0.97 at \(\sigma=21/40\). \(\mu(\sigma)\) is nearly flat at 0.87–0.95 across most of its range and then falls to 0.58 near \(\sigma=1\). \(\beta(\alpha)\) improves most at small \(\alpha\), reaching 0.25, and has a dead zone around \(\alpha \approx 0.30\) to \(0.45\) where three grid points never moved once. So within a single well-studied subfield, with no AI anywhere, the return to effort differs by a factor of five depending on which part of the parameter space is attacked. That is the incidence claim this project asks about, measured on human mathematics. The reading is the log’s.
  • No AI has contributed to any of these bounds, and the negative is sourced rather than assumed. Three separate things establish it. The database’s own authors say AI integration has not happened: “one could also imagine integrating the ANTEDB with other tools, such as Lean or AI systems, but for now we have focused primarily on collecting the data and optimizing the relations between the exponents,” and “the database only contains a placeholder Lean folder” [→ ANTEDB]. The automation that did produce new bounds is linear-programming-style optimization over collated relations, not a model, and produced them “without introducing any substantial new inputs from analytic number theory.” And when an AI system was pointed at this area, it failed: Tao reports AlphaEvolve “struggled to take advantage of the number theoretic structure in the problem, even when given suitable expert hints” [→ AlphaEvolve mathematics]. Every record change in every panel above is attributed to a human paper, the latest being Guth–Maynard 2024 and Trudgian–Yang 2023.
  • The stated reason is problem form, not difficulty, and it has a testable edge. Tao’s diagnosis distinguishes two possibilities and does not settle between them: “This could potentially be a prompting issue, or perhaps the landscape of number-theoretic optimization problems is less amenable to this sort of LLM-based evolutionary approach.” What did work needed algebraic structure a search could exploit — “AlphaEvolve does seem to do well when the constructions have some algebraic structure” — and these exponents are asymptotic inequalities rather than finite constructions with a computable score. That is a claim about the shape of the problem, so it predicts the AI-reachable part of mathematics is delimited by whether a candidate can be cheaply scored, not by how hard the mathematics is. Reading these panels alongside that failure is the log’s comparison, not Tao’s.
  • A caution about crediting the automation that did work. The exponent-database improvements are a good case for the argument that collation finds unpicked fruit, but the credit belongs to a human-designed relation database and a solver. The attribution audit of the AI-discovery literature makes the parallel point about AlphaEvolve’s mathematical results — that the search space and the domain knowledge in the prompt, rather than the evolutionary machinery, are what determine performance [→ simple baselines]. In both cases what did the work was a human’s formalization of where to look.
  • What this cannot support. The units are not comparable to a cost curve in the way the arithmetic above pretends. A cost curve’s denominator is money or compute; an exponent’s improvement has no denominator at all, because research effort per bound is unmeasured and certainly rose over the century — which is the fishing-out baseline that applies here as everywhere [→ ideas harder to find]. The grids are also uniform in the parameter, which is not a measure of mathematical interest: \(\sigma\) near 1 is easier territory as well as faster-moving, so the spread in the second figure mixes difficulty with attention. And none of these series has any AI in it: the latest entry is 2024 and the automated search that the database’s launch paper reports is optimization over collated relations rather than a model [→ ANTEDB]. So this fixes the pre-AI baseline for the math domain, which the domain previously lacked entirely, and measures no AI contribution whatever.
  • One year fails inside the database’s own solver, and is dropped rather than patched. The \(\beta\) sweep cannot compute 1991: compute_best_beta_bounds raises a TypeError for that restriction of the literature. The extraction script reports the skip rather than silently omitting it, and the 1991 \(\beta\) column is therefore missing from the sweep. No other year fails, and \(\mu\) and \(A\) are unaffected.
  • Dates: the underlying references run 1920 to 2024; extracted and plotted 2026-07-26 from the database as of that date.
  • Bears on: Q1 growth rate, Q6 intertemporal, Q7 incidence, Q8 benchmarks.
  • Links: extracted by tools/antedb_extract.py from the database recorded at ANTEDB; the CSVs are vendored at posts/data/apple-picking/antedb-bounds.csv and antedb-sweep.csv, and the figure is generated by tools/sources_figures.py.
  • Unused: held for the argument’s math efficiency-definition slug, which asserts that bounds give a dated continuous series without yet showing one, and for the Q1 baseline; not yet cited.
  • Status: derived — every value comes from the database entry named above via the extraction script, the halving times are this log’s arithmetic over those values, and nothing is quoted as though a source had said it. The two external checks are noted in the entry: the Guth–Maynard functional-form discrepancy, and the secondary account of the 1940–2024 plateau.

Inventory of the 67 AlphaEvolve problems, as the frame for a historical baseline (2026)

This log’s synthesis. Not a source: an inventory built here from the AlphaEvolve mathematics paper and its companion repository [→ AlphaEvolve mathematics], so that the historical baseline the known gaps ask for can be built on a stated frame rather than on whichever problems turned out to be convenient. It answers the prior question — which of these problems even has a record history to compare against — and extracts whatever history the paper already carries.

Two bar charts. Left: of 65 problems numbered in the paper, 19 are ones where AlphaEvolve holds the record, 11 matched a known optimum, 8 came in below the record, 4 have had their record surpassed since, and 23 are unclassified. Right: a descending sequence of bars from 65 numbered in the paper, to 31 with a live numeric record, to 12 sampled, to 6 that yielded a record sequence, to 2 with both AI and human steps.

Composition of the AlphaEvolve problem set, and how many problems survive each filter a historical comparison requires.

Left: the repository’s own status.json classification, applied to the 65 problems the paper numbers under the assumed index mapping described below. The 31 problems in the first, third and fourth categories are those with a live numeric record; the matched-optimal group has a terminated history and the unclassified group is mostly conjectures and non-record tasks. Right: the same set after each filter needed to compare an AI record step against the historical steps on the same quantity, ending at the two quantities where both exist [→ record steps]. Counts computed here from the two named inputs.

  • A prior-art search found nobody had done this. The nearest miss is HorizonMath, a benchmark of over 100 unsolved problems with automated verification, which compares AI output to best-known published results and carries no historical dimension [→ HorizonMath]. Dated record tables exist per problem — Packomania and Erich’s Packing Center for packings, Radziszowski’s Small Ramsey Numbers dynamic survey, the well-known tabulation of the matrix-multiplication exponent — and MathBases indexes roughly 400 such databases, but no one has joined them to the AI results. A general statistical literature on record progression exists for sports, biology and technology and does not cover mathematics. So the data and the method both exist and have not been put together.
  • The paper’s authors decline the historical survey explicitly, which is why it is missing. Their stated reason: “For reasons of space, we do not attempt to exhaustively survey the history of each of the problems listed here, and refer the reader to the references provided for each problem for a more in-depth discussion of known results.” The references are therefore the intended route to the history, and this inventory follows it.
  • How it was built. tools/alphaevolve_inventory.py reads a local pdftotext extraction of the paper plus a checkout of the companion repository, locates each problem’s definition inside the paper’s problem section, and records its title, topic group, the bracketed references cited within its span, the publication year of each of those references from the parsed bibliography, any inline bound string, and the repository’s status classification. Output is vendored at posts/data/apple-picking/alphaevolve-inventory.csv. Neither the paper nor its text is redistributed here.
  • The frame is 65 numbered problems, not 67, and roughly 50 distinct ones. The paper defines problems 6.1 to 6.65, while the repository’s status.json indexes 1 to 67, so the two enumerations cannot be identical and the identity mapping is an assumption every row records as such. Separately, at least twelve of the repository’s 67 experiment directories are the same problem under two names. Counted here.
  • Only about half the problems have a live record to compare against. Under the assumed mapping the statuses come out as 19 where AlphaEvolve holds the record, 11 where it matched a known optimum, 8 where it fell below the record, 4 where its result has since been surpassed, and 23 unclassified. The 31 in the first, third and fourth groups are the ones with a live numeric record; the matched-optimal group has a terminated history and the unclassified group is mostly conjectures and non-record tasks. So the sampling frame for a baseline is about 31 problems.
  • Most of the last one to three record steps are already dated, which was the surprise. Of the 65 problems, 63 cite at least one dated reference, 52 cite at least two, and 30 cite at least four; the bibliography yields 302 entries of which 298 carry a year. Spot-checked against problems whose histories are known independently: the Sidon autoconvolution problem returns 2010 and 2017, matching the Matolcsi–Vinuesa and Cloninger–Steinerberger attributions its own notebook gives, and the classic moving sofa returns 1992 and 2024, matching Gerver and Baek. Among problems citing two or more dated works the median span between earliest and latest cited year is 36 years, so these are decades-deep literatures rather than fresh ones.
  • What the dated citations are not. A cited year is the year of a cited work, not of a record improvement on that problem’s quantity: background references, surveys, and method papers are mixed in with the papers that moved the bound. Separating them requires reading the cited papers, which is the manual step this inventory scopes rather than performs. Nothing in the CSV should be read as a record sequence.
  • Two design constraints the inventory makes visible. First, the sample must be drawn from the 31-problem frame before the answers are looked at, because tractability correlates with being well-curated and so with progress rate — the selection warning this log quotes against itself applies with full force here [→ Grace]. Second, for the packing problems the historical baseline is itself computer search: record improvements have been machine-generated since the 1990s and many are unpublished, with Packomania reporting improvements arriving daily. On those problems an AI-versus-history comparison is automated-versus-automated, so whether each prior record was human-proved or machine-found has to be coded per problem or the headline comparison means nothing. Both points are this log’s.
  • Dates: the paper is arXiv 2025-11-03, and the extraction used v3 dated 2025-12-22; the repository was read at its state of 2026-07-26; the cited works span 1898 to 2025 and the inventory was built 2026-07-26.
  • Bears on: Q1 growth rate, Q7 incidence, Q8 benchmarks.
  • Links: built by tools/alphaevolve_inventory.py from arXiv 2511.02864 and the companion repository; the CSV is vendored at posts/data/apple-picking/alphaevolve-inventory.csv.
  • Unused: not yet cited in the argument, and it is a frame rather than a finding. It becomes citable when the record sequences are built on top of it.
  • Status: derived — every field is computed by the named script from the two named inputs and is reproducible from them; nothing is quoted as though a source had said it, apart from the authors’ own sentence declining the historical survey, which is quoted verbatim. Two extraction bugs were found and fixed during construction and are recorded because they would have corrupted the output silently: cross-references to a problem occurring before its definition caused the preceding problem’s span to swallow its content, and the topic-group headings were only partly matched so group labels drifted forward. Both were caught by checking two problems whose histories are known independently, which is the check to repeat if the script is changed.

AI record steps against human steps on the same quantities (2026)

This log’s synthesis. Not a source: dated record sequences built here for a pre-committed sample of twelve AlphaEvolve problems, drawn from the frame the inventory established [→ AlphaEvolve inventory], so the AI-era step can be set against the historical steps on the same quantity. This is the comparison the discussion of that paper does not contain.

Three panels. Left: the sums-and-differences lower bound rising over eight record steps from 1.079 to 1.173, four steps dated 2007 by humans, two dated 2025 by AlphaEvolve, and two dated 2025 by humans at the top. Middle: the autoconvolution constant over four steps from 0.889 in 2010 to 0.961 in 2025, alternating between AlphaEvolve and a human gradient-based method. Right: a strip plot of relative improvement per record step on a symmetric log axis, with AlphaEvolve at median plus 0.91 percent, human computer search at plus 2.52 percent, and human by-hand work at plus 1.26 percent.

Record sequences for the two sampled quantities with both AI and human steps, and the pooled distribution of step sizes by who made each step.

Left and middle: the best known value against record step, not against year, because several steps share a year. Marker colour is who made the step. Each point is annotated with its year. Right: every step in the sample with a computable gain, plotted as the relative improvement in the bound, on a symmetric log axis so that near-zero steps remain visible; the vertical bar is each group’s median. Values are transcribed from the paper’s prose and the agent coding is this log’s; the medians are its arithmetic.

  • How the sample was drawn, before any values were read. The frame is the 31 problems whose status is world_record, worse_than_record or former_record — those with a live numeric record. Twelve were selected as the smallest SHA-256 digests of a declared salt joined to the problem label, a rule anyone can recompute: 6.1, 6.3, 6.4, 6.9, 6.10, 6.30, 6.35, 6.36, 6.38, 6.40, 6.42, 6.44. The draw happens to contain all four former_record problems, four of twelve against four of thirty-one, so already-surpassed problems are over-represented about two and a half fold. It is kept rather than redrawn, because redrawing on inspection is the thing pre-commitment exists to prevent.
  • Half the sample has no scalar record sequence at all, and that is the main finding. Six of the twelve could not be reduced to a dated series of numbers: 6.1 and 6.9 and 6.10 state bounds asymptotically or as inequalities in a parameter rather than as a value; 6.38 is a table over a range of \(N\) improved “in several instances”; 6.36’s values live in the repository rather than the paper; and on 6.40 AlphaEvolve produced sequences shorter than the standing record, so there is no AI step to place. The same functions-not-series problem that the exponent-database work ran into [→ ANTEDB rates] recurs here and bites harder, because these problems have no equivalent of a solver to evaluate a slice.
  • Where a head-to-head is possible at all, it is two quantities. Only 6.3 and 6.44 carry both AI and human steps. Everywhere else the paper states the incumbent value and no prior sequence, so there is exactly one earlier number and no distribution to compare against. Anyone reading a stronger claim than “two cases” out of this table is reading too much.
  • In those two cases the AI steps are ordinary in size, and on one of them smaller than the human steps around them. On 6.44 the recorded steps are, in order, +2.65%, +0.79% and +2.52% by humans in 2007, then +0.28% and +0.91% by AlphaEvolve in 2025, then +1.27% by a human in 2025. Both AI steps are smaller than the median human step on the same quantity. On 6.3 the pattern reverses: AlphaEvolve’s second step is +6.59%, the largest in the sample. Pooled across all sixteen steps the medians are +0.91% for AlphaEvolve, +2.52% for human computer search and +1.26% for human work by hand. Computed here; with n=10 and n=6 no inference is warranted beyond “the same order of magnitude.”
  • A human took back the record on 6.44 within months, in a paper titled after doing so. The sequence ends with two human improvements to 1.173050 and 1.173077, above AlphaEvolve’s 1.1584, which the paper says were reached “by using mathematical methods closer to the original constructions of [158].” The first of them is Gerbicz, whose cited title is “Sums and differences of sets (improvement over AlphaEvolve).” So on the one sampled quantity with a real contest, insight-led human work overtook search-led AI work, and the paper records it.
  • On 6.3 the two approaches leapfrogged, and the AI’s larger step came after seeing the human’s. The order is: prior bound 0.88922; AlphaEvolve 0.8962 in “a quick experiment”; Boyer and Li independently 0.901564 by gradient methods; then, in the authors’ words, “Seeing this result, we ran our experiment for a bit longer,” reaching 0.961 with a 50,000-part step function. The largest step in the sample was therefore taken with knowledge of a competitor’s method, which is a different thing from an independent discovery.
  • Several steps are explicitly compute-bounded rather than idea-bounded, in the authors’ own words. On 6.3, “We believe that with even more parts, this lower bound can be further improved.” On 6.42, the construction “can likely be improved further.” On 6.35, it “can likely still be improved slightly by manual analysis.” And on the circle packing problem the authors generalize the point: “the problem allows for continued numerical refinement, where further gains are largely a function of computational investment.”7 On these quantities “who holds the record” is a statement about who last spent compute, not about who understands the problem best.
  • The provenance of the human baseline is often not a paper. The prior records cited for the packing problems are Erich Friedman’s records webpages; one of them is cited with literal [YEAR] and [DATE] placeholders left unfilled in the paper’s bibliography, so that step cannot be dated at all. Among the works that refined the circle packing construction, one is a commercial solver vendor’s blog post and another is a comment on a GitHub issue. Two further steps rest on a personal communication and on a reference that did not parse from the bibliography. So for a substantial share of these quantities the historical baseline is a community leaderboard maintained by continuous computer search, which is the confound the inventory flagged, now confirmed in the sample.
  • What would make this a measurement rather than an illustration. Full sequences, which means reading the cited papers rather than the citing one: the paper gives the incumbent and occasionally names one predecessor, and everything earlier has to come from the primary literature. On this sample that is roughly six problems’ worth of citation-chain walking. The honest summary of what twelve problems bought is a frame, two head-to-heads, and a strong result about why the exercise is harder than it looks.
  • Dates: the record steps span 1996 to 2025; the sample was drawn and the sequences transcribed 2026-07-26.
  • Bears on: Q1 growth rate, Q4 expertise, Q5 returns, Q7 incidence.
  • Links: built by tools/alphaevolve_records.py; values transcribed from arXiv 2511.02864 and recorded per step with their reference numbers; the CSV is vendored at posts/data/apple-picking/alphaevolve-records.csv.
  • Unused: not yet cited in the argument. It is the first direct evidence in the log on whether an AI record step is large or small by the standards of its own problem, and belongs wherever the argument discusses the depth of AI mathematical contributions.
  • Status: derived — every value is transcribed by hand from the paper’s prose, with the quoted sentence and the reference number recorded against each step in the script, so each is checkable against the source; the agent coding, the relative gains and the medians are this log’s. Nothing is quoted as though a source had said something it did not. Two limits are structural rather than incidental: the sample is twelve problems and the head-to-head is two quantities, and the paper is both the source of the AI values and the source of the historical ones, so a step it failed to mention is invisible here.

7 The full passage is worth having, because it is a vendor describing a race with no ceiling: “In our initial work, AlphaEvolve found new constructions improving these bounds. To adhere to the three-digit precision established in [129, 128], our publication presented a simplified construction with truncated values, sufficient to secure an improvement in the third decimal place. Subsequent work [25, 94] has since refined our published construction, extending its numerical precision in the later decimal places. As this demonstrates, the problem allows for continued numerical refinement, where further gains are largely a function of computational investment. A brief subsequent experiment with AlphaEvolve readily produced a new construction that surpasses these recent bounds.”

Every rate in the log that has a denominator (2026)

This log’s synthesis. Not a source. The log’s first standing rule is to prefer a denominator to a headline, and this is the inventory of where one exists. It is the fastest way to see that the yields cluster low and that the highest-looking figures are the ones measured on the most artificial task.

  • How it was built. Every entry above was read for a claim of the form “N of M” or an explicit percentage with a stated base; those are collected here with their anchors. Figures without a base are excluded by construction, which is why several of the log’s most-quoted results do not appear. Nothing here is quoted; each figure lives in its own entry with the source’s wording.
  • Open research problems, formal proof. 9 of 353 open Erdős problems, about 2.5%, and 44 of 492 OEIS conjectures, about 9% [→ AlphaProof Nexus]. Of roughly 47 AI-standalone Erdős contributions, about 13 full resolutions and about 9 incorrect [→ Erdős wiki].
  • Competition mathematics. 25 of 30 olympiad geometry problems, against 25.9 for the average gold medalist [→ AlphaGeometry]; 84% of 25 years of geometry problems, up from 54% [→ AlphaGeometry 2]; 5 of 6 IMO 2025 problems as a vendor claim [→ Aristotle]; 5 of 6 for the graded gold [→ IMO 2025].
  • Vulnerability discovery and exploitation. 54 of 63 planted vulnerabilities found and 43 patched, with 18 real zero-days and 11 patches [→ AIxCC]; about 132 confirmed-and-resolved of about 1,060 submissions, with about 208 duplicates [→ XBOW]; about 20% reproduction on 1,507 vulnerabilities [→ CyberGym]; 13 of 15 one-day CVEs with the description and about 1 of 15 without [→ Fang one-day]; 9 valid vulnerabilities at an 82% valid-submission rate, second of eleven [→ ARTEMIS]; 15.6 of 32 attack steps at the highest budget, and 1.2–1.4 of 7 on the industrial range [→ AISI cyber].
  • Software and ML engineering. 1.96% of 2,294 GitHub issues at launch [→ SWE-bench]; bronze-medal level in 16.9% of 75 Kaggle competitions [→ MLE-bench]; a 21.0% average replication score on 20 ICML papers [→ PaperBench]; 21% on the hardest of 270 reproducibility tasks [→ CORE-Bench]; under 5% success on 102 optimization tasks [→ GSO]; 0.23× of expert speedup on 498 tasks [→ SWE-fficiency]; 1 of 3 autonomous manuscripts above a workshop threshold [→ AI Scientist-v2]; 47.6% wins-or-ties on a 220-task subset [→ GDPval].
  • The pattern, which is this log’s reading and not any source’s. Where the task is a fixed set of real problems and the denominator is published, the yield is almost always under a quarter, and often under a tenth. The exceptions are competition mathematics, where problems are constructed to be solvable in hours, and the one-day exploitation case, where the human-written description supplies the localization. Both exceptions are informative about what makes a task tractable rather than counterexamples to the pattern.
  • What the inventory cannot do. These denominators are not commensurable. A planted synthetic bug, an open Erdős problem, and a Kaggle competition differ in difficulty by an unknown amount, so the figures cannot be averaged or ranked across rows. The inventory establishes the order of magnitude of published yields and the fact that most entries have no denominator at all, nothing finer.
  • Dates: assembled 2026-07-26 from entries dated 2016 to 2026-07.
  • Bears on: Q1 growth rate, Q2 autonomy, Q5 returns, Q7 incidence, Q8 benchmarks.
  • Links: every input is an entry in this document; the external links are on those entries. Assembled by hand rather than by script, so a reader re-checking it should re-read the named entries (validator checks anchors resolve, not arithmetic).
  • Unused: held for the argument’s Q5 and Q8 discussion; not yet cited.
  • Status: derived — every figure is quoted in the entry it comes from, and this entry quotes nothing itself.

Every cost-per-result figure in the log (2026)

This log’s synthesis. Not a source. Cost per result is the most-quoted number in this literature and the least comparable, and Bloom and co-authors argue it is the wrong measure in principle [→ ideas harder to find]. This entry collects the figures so the range is visible and the objection can be applied to all of them at once.

  • How it was built. Every dollar-denominated figure in the log, with what it is denominated per. Vendor figures are marked, because most of them are vendor figures.
  • Per confirmed vulnerability or exploit. About $152 per competition task [→ AIxCC]; $12,500 per attempt at a 100M-token budget, $125k for ten runs [→ AISI runs]; a 27-year-old OpenBSD bug at under $20,000 across about 1,000 scaffold runs, with the single successful run under $50, and FFmpeg vulnerabilities at roughly ten thousand dollars over several hundred runs, all vendor-reported [→ Mythos]; a full root exploit from a known vulnerability at under $1,000 and half a day, vendor-reported [→ Mythos]; about $42,000 per bug implied by a contested deployment report [→ Palo Alto]; detection of a showcased overflow by a model priced at $0.11 per million tokens [→ AISLE].
  • Per hour of work. $18 an hour for some agent variants against $60 an hour for professional penetration testers [→ ARTEMIS].
  • Per mathematical result. A few hundred dollars per resolved open Erdős problem, with the full system saving 2× to 5× over a basic verify-and-retry loop on the hardest cases [→ AlphaProof Nexus].
  • Per algorithmic or research artifact. A hard $1 per task cap, under which only surface-level optimizations appeared [→ AlgoTune]; a few hundred dollars of test-time compute for state-of-the-art kernels, with the accounting unresolved [→ TTT-Discover]; under $15 per generated paper, judged by the authors’ own automated reviewer [→ AI Scientist]; $50,000 in model credits per team from each of three labs for the AIxCC final [→ AIxCC].
  • The range is about six orders of magnitude, and that is the finding. From $1 per optimization task to about $42,000 per bug. This log’s reading: the spread is set almost entirely by the difficulty and realism of the target, not by the price of tokens, so a cost-per-result figure is a statement about the task and only incidentally about the system. Two figures from the same vendor post differ by more than two orders of magnitude depending on whether failed runs are counted [→ Mythos], which is the clearest single demonstration.
  • Why none of these measures returns, stated by the source that says it best. Bloom, Jones, Van Reenen, and Webb: ideas per research dollar is predicted to decline by essentially every idea-driven growth model, so “these natural measures are not really informative about whether research faces constant or diminishing returns” [→ ideas harder to find]. The theory-relevant object is ideas per researcher. Every figure above is the disqualified measure.
  • What a usable version would need. All attempts counted, not only successful ones; a fixed target population; and the human cost of choosing the target and verifying the output included. One entry approaches this by reporting cost against solve rate at a fixed problem [→ AlphaProof Nexus]; nothing else does.
  • Dates: assembled 2026-07-26 from entries dated 2020 to 2026-07.
  • Bears on: Q1 growth rate, Q5 returns, Q6 intertemporal, Q8 benchmarks.
  • Links: every input is an entry in this document, and the external links are on those entries. Assembled by hand; the source log is the only reference needed to re-check it.
  • Unused: held for the argument’s Q5 discussion; not yet cited.
  • Status: derived — every figure is quoted in the entry it comes from, and this entry quotes nothing itself.

What “autonomous” turned out to mean, case by case (2026)

This log’s synthesis. Not a source. Q2 asks whether AI can produce a complete result with no human driving it, and nearly every entry claiming so means something different by it. This is a coded ladder, applied to the log’s autonomy claims, so that the question can be asked at a fixed rung instead of re-litigated per case.

  • How it was built. Five rungs, defined below, then each of the log’s autonomy claims assigned to the highest rung its own entry supports. The assignment is this log’s judgment about what the entries say, not a claim by any source, and a reader who reads an entry differently should move it.
  • Rung 1, execution on a specified subproblem. The human states the problem, supplies the localization, and verifies. The exploitation results with the CVE description in hand are here, and the twelvefold drop when the description is withheld is the measurement of how much rung 1 was contributing [→ Fang one-day]. Benchmark scores are almost all rung 1 by construction, because a benchmark supplies the target [→ Naptime for the authors saying so].
  • Rung 2, search within a human-built harness on a human-chosen target. The model does not know the answer, but an expert built the scaffold and the verifier. Most of the log’s strongest results sit here: the kernel records [→ TTT-Discover], the evolutionary coding results [→ AlphaEvolve, FunSearch], the formal proof searches [→ AlphaProof Nexus, Aristotle], the leaderboard records [→ nanogpt], and the pre-LLM systems that show this rung did not require language models [→ Cyber Grand Challenge, AlphaTensor and AlphaDev].
  • Rung 3, unaided discovery on a target class, with human verification after. The system finds something nobody had specified, and an expert confirms it. Big Sleep’s twenty open-source bugs, with a human only in final review, and the live SQLite zero-day are here [→ Big Sleep], as are the AIxCC finalists’ 18 real zero-days [→ AIxCC] and CyberGym’s incidental 34 [→ CyberGym].
  • Rung 4, a complete result the field accepts, on a problem the field cared about. The unit-distance disproof is the log’s clearest case, and it is unusual in that the model was reportedly not specialized, not scaffolded for proof search, and not aimed at the problem [→ unit-distance]. Erdős #728 is a weaker instance, hedged by an operator-in-the-loop convention [→ Erdős 728]. One autonomously generated manuscript above a workshop acceptance threshold is a rung-4 claim about a much lower bar [→ AI Scientist-v2].
  • Rung 5, choosing what to work on. No entry in this log reaches it. Every case above has a human selecting the target, the problem class, or the corpus. The log records no instance of a system choosing a research agenda and being judged to have chosen well.
  • What the ladder shows, and it is the log’s reading. The autonomy claims cluster at rung 2, the headline claims that travel furthest are rung 3 and 4, and the gap between rungs 2 and 4 is almost entirely a question of who supplied the verifier and who chose the target — not of what the model did inside the loop. That is why the argument’s closing caveat, that autonomous almost always means autonomous execution, is the right reading, and it is also why “autonomous” in a vendor headline cannot be compared across entries without doing this coding first.
  • Where the coding is contestable. Rung 4 for the unit-distance result rests on the vendor’s account of what the model was and was not given, which no outside party can check [→ unit-distance]. The rung-3 cases rest on how much the human final review contributed, which is nowhere itemized. Both are the log’s assignments and both could move a rung on better information.
  • Dates: assembled 2026-07-26 from entries dated 2016 to 2026-05.
  • Bears on: Q2 autonomy, Q4 expertise, Q8 benchmarks.
  • Links: every input is an entry in this document, and the external links are on those entries. Nothing here is quoted from a source; the rungs are this log’s construction.
  • Unused: held for the argument’s Q2 discussion, which currently treats autonomy as a threshold rather than a ladder; not yet cited.
  • Status: derived — an assignment of the log’s own entries to categories the log invented, reproducible from those entries and quoting nothing.

The randomized and quasi-experimental evidence, in one place (2026)

This log’s synthesis. Not a source. The log’s evidence ranking puts randomization at the top, and only a handful of entries qualify. Collected here because they disagree about sign, and because the disagreement is structured rather than noisy — which is more useful than any one of them.

  • How it was built. Every entry whose design randomizes AI access or exploits a plausibly exogenous rollout, with its sign, setting, and sample. Effect sizes are quoted in the entries, not here.
  • Randomized, positive, on constructed tasks. Writing tasks, about 450 professionals, large positive with compression toward lower-ability workers [→ Noy and Zhang]. An HTTP-server implementation, about 35 completers per arm, 55.8% faster with a 21–89% interval and a null on success rate [→ Copilot RCT]. A business problem-solving exercise, 1,174 adults, about three quarters of the education gap closed [→ education gap]. Consulting tasks, 758 consultants, positive inside the frontier and 19% worse outside it [→ jagged frontier].
  • Randomized, negative or null, on real work. Real issues on the contributors’ own mature repositories, 16 developers and 246 issues, 19% slower [→ METR RCT]. A security-related programming task, 159 developers, no significant effect on code security and experience not substitutable [→ Gemini and developer experience]. Kenyan small-business owners, 640 firms, null on average and negative for initial low performers [→ Kenya].
  • Randomized, positive, pooled across firms. Three trials at three firms, 4,867 developers, about 26% more completed tasks — and this log has not verified it against the working paper [→ pooled RCTs].
  • Quasi-experimental, on rollouts and thresholds. Staggered deployment to about 5,000 support agents, 14% average and 34% for novices [→ support agents]. A Copilot eligibility discontinuity, 187,489 developers, composition shifted toward coding and away from project management [→ Copilot composition]. Freelance platform difference-in-differences, 92,547 freelancers, employment and compensation down and top freelancers hit hardest [→ freelancer demand]. Job postings, automation-prone categories down 21% [→ posting demand]. Payroll records, entry-level employment in exposed occupations down 16% [→ canaries]. French firms 2017–2020, employment and sales up [→ French firms].
  • The structure of the disagreement, which is this log’s reading. Sign tracks two things. It tracks the quality of the starting point: gains are positive on constructed or unoptimized tasks and zero-to-negative on mature repositories with high standards, which is the same collapse the optimization benchmarks show [→ SWE-fficiency, GSO]. And it tracks whether the task has a checkable answer: where it does, AI compresses the skill distribution; where the user must judge which output to trust, it widens it. Both patterns are established by randomization, in different studies, and neither is established within a research domain.
  • What the whole set cannot do. Every randomized entry measures task execution, none measures discovery, and none is in cyber, math, or algorithms except one underpowered null on code security [→ Gemini and developer experience]. So the strongest designs in the log are the furthest from its subject, and the entries closest to its subject are demonstrations. That trade-off is the central measurement problem of this document, and no entry escapes it.
  • Dates: assembled 2026-07-26 from entries whose fieldwork runs 2020 to 2026.
  • Bears on: Q1 growth rate, Q3 demand, Q4 expertise, Q7 incidence.
  • Links: every input is an entry in this document, and the external links are on those entries. Assembled by hand from the entries’ own design descriptions; see the evidence-weighting rules above for how the tiers were defined.
  • Unused: held for the argument’s Q4 and Q7 discussion; not yet cited.
  • Status: derived — a classification of the log’s own entries by design, quoting nothing and reproducible from them.

Known gaps

Where the evidence is thin, by question. These are the entries a future version of this log most needs, and their absence is the main reason the argument’s verdict is a ranking of theories rather than a measurement. Each paragraph says what has been added since the gap was first written, so the section records progress rather than restating the same complaint.

Four stacked panels sharing a year axis from 1915 to 2032. Panel one, analytic number theory exponents: thirty rows, each with dots at the years its record moved, with mu rows starting in 1920, A rows in 1921 and beta rows only around 1989. Panel two, twelve AlphaEvolve problems: six rows are dotted with a cross and a stated reason for having no scalar record, six carry dots clustered at 2025, and two undatable steps sit in a shaded date-unknown gutter at the right. Panel three, sixty-six physical technology cost curves as horizontal spans, all ending by 2013. Panel four, three AI algorithmic-efficiency estimates shown only as measurement windows, with a note that no per-year series is published.

What dated evidence exists in each of the log’s four independent problem sets, with missing data drawn rather than omitted.

The four problem sets this log draws efficiency evidence from, each on its own panel and never pooled, on one shared year axis. Rows are individual problems. A filled dot is a dated improvement, a solid line a stretch over which the value is known and unchanged, an open dot an improvement whose date could not be established, a dotted line a problem with no dated series available, and a cross a problem in the set for which no scalar record exists. Set 1 is the exponent database [→ ANTEDB rates], set 2 the sampled AlphaEvolve problems [→ record steps], set 3 the technology cost curves [→ OWID], and set 4 the AI algorithmic-efficiency estimates [→ Epoch on LMs, ImageNet, compute-to-AlexNet]. Sets 1 and 3 contain no AI. The symbols and the grouping are this log’s; the dates are as recorded in the entries named.

Read across the panels, two things are visible that no single entry states. The sets barely overlap in time, so comparisons between them are across eras rather than like-for-like. And the density of dated evidence is very uneven: set 1 holds hundreds of dated improvements over a century, while set 2 holds sixteen steps, half its problems empty, and two undatable.

Q1 growth rate is measured only in proxies, and the proxies now have a good baseline. No source here estimates AI’s effect on a domain-level efficiency curve of the kind the OWID series plot. The pretraining curve is measured but ends before agents mattered [→ algorithmic progress]; the RCT measures a task-level effect in one setting [→ METR RCT]. Nothing connects the two. What has improved is the counterfactual rather than the estimate: the log now holds three independent pre-AI algorithmic-progress series measured with hardware physically held constant or removed [→ Sherry and Thompson, Bixby, SAT Museum], the hardware curve they should be compared against [→ hardware price-performance], and two macro estimates whose authors both state that the ideas channel is excluded from them [→ Acemoglu, Aghion and Bunel]. The math domain, which had no efficiency curve of any kind, now has six, extracted from the exponent database and running 1920 to 2024 [→ ANTEDB rates]. So the rate an AI contribution would have to beat is now well characterized in all three domains, and the contribution itself still is not measured in any of them.

Nobody has put the AI mathematics results against a historical baseline, and this log’s exponent series is the method that would. The AlphaEvolve mathematics paper improved bounds on roughly a fifth of 67 problems [→ AlphaEvolve mathematics], and the discussion of it splits into two unsatisfying halves. One half is rhetorical — “decades of human effort,” “untouched for over 50 years” — which is a claim about a baseline without a baseline. The other quotes magnitudes, mostly fourth- and fifth-decimal-place nudges, and defends them qualitatively on the grounds that in these areas “progress can be at times glacial.” Neither assembles what would settle it: the dated record of prior improvements on those same problems, so the AI-era step can be compared with the distribution of historical step sizes and the intervals between them. Two per-problem anchors exist in the coverage — 56 years for Strassen, and an Erdős minimum-overlap record that “hadn’t budged since 2016” — and they differ by a factor of six, which is why anecdotes cannot substitute for the distribution. That comparison is exactly what this log built for the analytic-number-theory exponents [→ ANTEDB rates]. Doing the same for the 67 problems is the single highest-value addition to the math domain, and it is tractable: the problems are named and their literatures are dated. Until someone does it, “AI improved 20% of these bounds” and “human mathematics improves these bounds all the time” are both true and neither is informative.

In math the gap is not thin evidence but a sourced absence, which is a different kind of finding. For cyber and algorithms the problem is that nobody has measured AI’s effect on a domain efficiency curve. For the exponent bounds the position is stronger and stranger: there is now a century-long curve, and no AI has moved any part of it. The database’s authors describe AI integration as a future possibility they have not pursued, the automation that did produce new bounds is a solver over collated relations rather than a model, and the one recorded attempt to point an AI system at analytic number theory failed even with expert hints [→ ANTEDB rates, ANTEDB, AlphaEvolve mathematics]. So the math domain supplies a clean pre-AI baseline and a clean null, and the argument should not treat its Q1 row as merely unmeasured there. What is genuinely missing is the other side: a dated attempt, at known cost, on a fixed set of these exponents, which would turn the null into a measurement.

The three domains’ baselines differ by three orders of magnitude, which reframes what “bending the curve” would mean. Pretraining efficiency halves in about 8 months and the fastest physical cost curve in about 8.6 [→ Epoch on LMs, efficiency rates]; classical algorithmic progress moves in jumps every 3 to 5 years [→ SAT Museum]; analytic number theory’s exponents halve on timescales of 82 to 1,204 years [→ ANTEDB rates]. An AI contribution that would be invisible against the pretraining curve would be transformative against the exponent curves, so a single question about whether AI raises “the rate of efficiency growth” is really three questions with different answers, and the argument’s Q1 row does not currently separate them. This is the log’s reading.

Q3 demand has quasi-experimental estimates now, but none of them is in a research occupation. curl’s program closure and ARTEMIS’s hourly rates measure neither headcount nor wages [→ curl, ARTEMIS], and the cyber labour indicators come from an industry survey and a job-posting scrape with no counterfactual [→ cyber labour]. The additions since do have counterfactuals and point in different directions: freelance employment and compensation down with top freelancers hit hardest [→ freelancer demand], postings for automation-prone work down 21% [→ posting demand], entry-level employment in exposed occupations down 16% [→ canaries], and French firm-level employment up in a pre-generative period [→ French firms]. Two of those cover software developers, which is the closest any of it comes to this project’s domains. Occupation-level causal evidence for security researchers, mathematicians, or ML researchers specifically remains absent, and the entry-level concentration is the finding most worth extending to them, since it is the junior work through which the next generation of researchers trains.

Q4 expertise has been rebuilt out of domain, and the sign is contested. The most-cited result on how AI interacts with researcher ability has been disavowed [→ Toner-Rodgers]. The randomized literature added since points both ways, and the disagreement is structured rather than noisy. Where the task has a checkable right answer, AI compresses the skill distribution: support agents gained 34% at the novice end and nothing at the top [→ support agents], and an online experiment closed three quarters of the education gap [→ education gap]. Where the task is open-ended and the user must judge which suggestion to act on, AI widens it: Kenyan entrepreneurs split +15% against −10% by baseline ability, through selection among suggestions rather than differences in the suggestions themselves [→ Kenya]. Research is the open-ended case, and the METR developer RCT’s mechanism — uncritical acceptance of output being the cost — is the same one [→ METR RCT].

Two things are still missing. None of this is measured *in* cyber, math, or algorithms, except one underpowered null on code security [→ [Gemini and developer experience](#src-gemini-dev-security)]. And none of it measures discovery, as against execution of a defined task. A study measuring discovery outcomes by user expertise in one of these three domains remains the highest-value single addition.

Q5 returns lack a repeated-run protocol. Duplicates and crossovers are inferred from operational data rather than designed experiments [→ XBOW, RE-Bench]. Nobody has run the same agent at equal budget repeatedly against a fixed target population and reported the yield curve. The nearest thing is AlphaProof Nexus’s agent ablation, where a basic verify-and-retry loop reached the same nine problems as the full system at 2–5× the cost on the hardest ones [→ AlphaProof Nexus] — which is a scaffold comparison at fixed problem set, not a yield curve, but it is the closest available test of whether more machinery extends reach or only cuts cost.

A baseline for “returns diminish” is now in the log, and it is not zero. Research productivity was falling by roughly 5% a year across the whole economy long before AI, halving about every 13 years, and by 6.8% a year in semiconductors [→ ideas harder to find]. Any apple-picking prediction of falling yield has to beat that baseline to be distinctive. No entry here makes that comparison.

Q6 intertemporal has one clean series and otherwise almost no direct evidence. The staircase question needs a dated capability series at fixed scaffold. AIxCC comes closest, because the organizers held the competition structure fixed across the August 2024 semifinal and the August 2025 final [→ AIxCC] — though the teams rebuilt their systems between the two, so even that measures the stack rather than the models. AlphaGeometry’s 54% to 84% on a fixed 25-year problem set is a second [→ AlphaGeometry 2], with the same defect: model and scaffold moved together. Every other available series confounds model generation with harness, prompting, and task mix [→ PERFOPT, TTT-Discover, FrontierMath], and the size of that confound has now been measured directly: harness changes alone moved a cyber benchmark by up to twentyfold at fixed model [→ Naptime]. This remains the weakest-supported row in the argument’s table.

The staircase, if it appears, would not be diagnostic of AI. The one domain in this log with three decades of dated, hardware-controlled progress measurement shows exactly the pattern apple-picking predicts — slow years punctuated by jumps “with a frequency of 3 to 5 years” — with no AI involved at any point [→ SAT Museum]. Sherry and Thompson chart the same step functions across 113 algorithm families [→ Sherry and Thompson], and Bixby’s version-to-version speedups are lumpy in the same way inside one product line [→ Bixby]. So finding a staircase in an AI capability series would not distinguish AI-driven progress from ordinary algorithmic progress, and the argument’s decision to demote the staircase from a test to a question about release dynamics is if anything understated. This comparison is the log’s.

Dating is itself uneven, and the gaps are informative. Two entries cannot be pinned to a month: DeepMind’s validation-bottleneck essay carries no posting date, and XBOW’s own post carries none either [→ validation bottleneck, XBOW]. The Mythos preview’s page metadata post-dates a response to it, so its true posting date is unresolved [→ Mythos]. In all three cases the undated source is a vendor or advocacy document, and in all three the date would bear on how much independent work a competing claim could represent.

Q7 incidence has no controlled starting-point test. The pre/post comparison runs across different benchmarks rather than across staged versions of one codebase [→ AlgoTune, SWE-fficiency]. A matched experiment holding model, scaffold, budget, and metric fixed would settle it.

Cross-cutting: failure denominators are usually missing. Most entries report successes without the attempts that produced them. Where a denominator exists it is recorded above and collected in one place [→ denominators]; where it does not, the figure cannot support a rate. The inventory shows that where a real problem set does have a published denominator, the yield is almost always under a quarter — so the missing denominators are unlikely to be missing at random.

Cross-cutting: the strongest designs are the furthest from the subject. Every randomized entry in the log measures task execution rather than discovery, and only one of them is in cyber, math, or algorithms — an underpowered null on code security [→ Gemini and developer experience]. Everything in the three domains is a demonstration, a benchmark, or an operational log. The single highest-value addition to this document would be a randomized or staged experiment measuring discovery outcomes in one of the three domains, by user expertise. Nothing here is a substitute for it, and the syntheses assembled above are an attempt to get as far as possible without it [→ experimental evidence].

Cross-cutting: benchmark validity is now better documented than benchmark performance. The log holds a general critique of agent benchmarking [→ agents that matter], a contamination result on the most-quoted coding benchmark [→ SWE-bench illusion], a machine-and-scoring-rule audit of the optimization benchmarks [→ benchmark reliability], a vendor conceding its own kernel harness was gameable [→ robust-kbench], a funding-disclosure failure on the headline math benchmark [→ FrontierMath], and the clearest available case of a proxy being optimized while the objective did not move [→ AI-discovered drugs]. Taken together these do not show that measured progress is illusory; they show that no single benchmark score in this document should be load-bearing on its own. That is a conclusion about method, and it is the log’s.

Chronology

Every dated event in the log in a single order, so that claims about sequence and rate can be checked without reading the entries. Rows covering a span rather than a moment — the cost-curve and algorithmic-progress measurement windows — are placed at the year the span opens. Publication dates are used where the underlying work is not separately dated; arXiv revisions are listed only where a figure could have moved. This table is derived from the entries’ **Dates:** lines and must be updated with them, and it sorts strictly by start date, so a new row goes in position rather than at the end.

Scatter plot with eight rows, one per section of the log, and a horizontal axis of years from 2016 to 2027. Each point is a dated event, with a count per row on the right. Points are sparse before 2024 and dense across 2025 and 2026, with cyber, math and algorithms carrying the most events.

Every dated event in the log, by the section its entry sits in.

Generated by tools/sources_figures.py from the table below, so it cannot drift from it. One point per dated row, placed on the row of the section its entry belongs to, with the number of events per section at the right. The concentration in 2025 and 2026 is a property of the evidence base rather than of the log’s coverage: most of the primary sources on AI contributions in these domains were published in those two years.

Date Event Entry
1920–2024 Span of the six analytic-number-theory exponent series
1929–2013 Span of the 66 technology cost curves
1940–2019 Span of the 113-algorithm-family improvement survey
1946 Erdős poses the unit-distance conjecture
1988–2004 Span of the LP solver speedup measurement
1990s–2022 Span of the SAT Museum’s thirty years of solvers
2006–2023 Span of the ML hardware price-performance measurement
2012–2019 Span of the compute-to-AlexNet efficiency measurement
2012–2023 Span of the language-model algorithmic-progress data
2012 Bixby: MIP solvers 29,000× faster from algorithms alone; LP progress stopped after 2004
2013-08-03 Grace releases Algorithmic Progress in Six Domains; algorithms worth 50–100% of hardware
2013-12-09 Grace’s report last revised
2016-08-04 DARPA Cyber Grand Challenge: autonomous find-and-patch, nine years before AIxCC
2017-09-08 Bloom, Jones, Van Reenen, and Webb circulate Are Ideas Getting Harder to Find?
2017-10 Aghion, Jones, and Jones put AI in the idea production function
2019 GPT-2, start of the offensive-cyber horizon series
2020-04 Are Ideas Getting Harder to Find? published in the AER
2020-05-08 Hernandez and Brown define algorithmic progress as compute-to-past-capability; 44× since 2012
2021-09-20 Sherry and Thompson: half of 113 algorithm families show little or no improvement, 14% transformative
2021-11 How Fast Do Algorithms Improve? appears in print in Proceedings of the IEEE
2022-10-05 AlphaTensor beats a fifty-year matrix-multiplication record
2022-12-10 Erdil and Besiroglu: ImageNet compute requirements halve every nine months
2022-12-15 Besiroglu, Emery-Xu, and Thompson: AI R&D is more capital-intensive
2023 The SAT Museum re-runs thirty years of solvers on one machine
2023-02-13 Peng and co-authors: Copilot RCT, 55.8% faster on an HTTP-server task
2023-03-02 Noy and Zhang: ChatGPT compresses the writing productivity distribution
2023-04-20 Brynjolfsson, Li, and Raymond: support agents, +14% overall, +34% for novices
2023-06-07 AlphaDev’s sorting routines enter the C++ standard library
2023-06-09 Trudgian and Yang post the precursor exponent tables
2023-09 BCG jagged-frontier experiment circulated: +12.2% inside, −19% outside
2023-10-10 SWE-bench: the best model resolves 1.96% of 2,294 real GitHub issues
2023-11-09 Epoch: ML hardware price-performance doubles every 2.1 years
2023-12-09 Ide and Talamas: autonomous AI helps the knowledgeable, assistive AI the least
2023-12-14 FunSearch: first LLM result on an open problem, cap sets and bin packing
2023-12 Hui, Reshef, and Zhou: freelance jobs down 2%, compensation down 5.2%
2024 Otis and co-authors run the Kenyan entrepreneur RCT: +15% high, −10% low
2024-01-17 AlphaGeometry solves 25 of 30 olympiad geometry problems
2024-03-09 Algorithmic progress in language models, arXiv
2024-03-17 Korinek and Suh: wages collapse only if human task complexity is bounded
2024-04-05 Acemoglu: no more than 0.71% TFP over ten years, and less if tasks are hard to learn
2024-04-11 Fang and co-authors: 87% of one-day CVEs exploited with the description, 7% without
2024-05-28 modded-nanogpt baseline, 45 min
2024-06 Aghion and Bunel: 0.68pp median, with the ideas channel explicitly excluded
2024-06-02 Fang and co-authors: teams of agents on zero-day vulnerabilities
2024-06-20 Project Naptime: security tooling moves a benchmark up to twentyfold at fixed model
2024-07 AlphaProof and AlphaGeometry 2 reach IMO silver with multi-day compute
2024-07 Doshi and Hauser: AI raises individual creativity, cuts collective diversity ~10%
2024-07-01 Kapoor and co-authors: AI agents that matter
2024-07-14 Noy and Zhang published in Science, with different figures
2024-07-15 PutnamBench: 1,692 formalizations, solvers clear “a handful”
2024-08 AIxCC semifinal: 37% of synthetic bugs found, 25% patched
2024-08 GPT-4o, start of the AISI cyber-range series
2024-08-02 Meta CYBERSECEVAL 3: Llama 3 fails every stage past reconnaissance
2024-08-12 The AI Scientist: automated papers at under $15 each
2024-08-15 Cybench: agents clear tasks humans solved in up to 11 minutes
2024-09-17 CORE-Bench: 21% on the hardest reproducibility tasks
2024-09-25 The Equational Theories Project launches
2024-10-02 Song and co-authors: Copilot raises contributions 5.9%, coordination time 8%
2024-10-09 MLE-bench: bronze-medal level in 16.9% of 75 Kaggle competitions
2024-10-27 Hoffmann and co-authors: Copilot shifts developers from coordination to coding
2024-11 Big Sleep’s first SQLite find, in a development branch
2024-11 Toner-Rodgers preprint appears
2024-11-22 RE-Bench, arXiv
2024-12-10 Hao and co-authors: AI expands individual impact, contracts collective focus
2024-12-20 Epoch discloses OpenAI’s funding of and access to FrontierMath
2024-12 Demirci, Hannane, and Zhu: postings for automation-prone work down 21%
2025-01-28 Tao launches ANTEDB with automated exponent-pair improvements
2025-02–06 Window of the METR developer RCT
2025-02-05 AlphaGeometry 2: 84% of 25 years of geometry problems, up from 54%
2025-02-20 MLGym, arXiv
2025-03-12 Epoch: inference prices fall 9× to 900× a year depending on the milestone
2025-03-18 METR time horizons: 50% horizon doubling every ~7 months since 2019
2025-04-02 PaperBench: a 21.0% replication score, below the human baseline
2025-04-10 The AI Scientist-v2: one of three manuscripts above a workshop threshold
2025-05 AlphaEvolve raises the 11-dimensional kissing bound 592→593
2025-05 Aghion and co-authors: French firms adopting AI grew employment and sales
2025-05-09 Cui and co-authors pool three coding-assistant RCTs: +26% tasks, 4,867 developers
2025-05-14 AlphaEvolve announced; 48-multiplication matrix result
2025-05-16 MIT disavows Toner-Rodgers
2025-05-29 GSO, arXiv
2025-06 XBOW tops a HackerOne leaderboard
2025-06-03 CyberGym, arXiv
2025-06-14 The SWE-bench illusion: 76% bug localization without the repository
2025-07 Gemini Deep Think officially graded IMO gold
2025-07-10 METR RCT: AI made experienced developers 19% slower
2025-07-15 Big Sleep’s live SQLite zero-day, CVE-2025-6965
2025-07-19 AlgoTune, arXiv
2025-07-28 Harmonic announces Aristotle’s formally verified IMO 2025 proofs
2025-08-04 Big Sleep discloses 20 novel open-source bugs
2025-08-08 AIxCC final: 86% found, 68% patched, 18 real zero-days, ~$152/task
2025-08-18 Zhao and co-authors: AlphaFold raised cross-disciplinary collaboration 0.48%
2025-09-11 First AI-set modded-nanogpt record (hiverge.ai, 2.625 min)
2025-09-16 Sakana concedes kernel benchmarks have exploitable loopholes
2025-09-23 DORA 2025: throughput sign flips positive, stability still negative
2025-10 GPT-5 “solves” Erdős problems — actually literature retrieval
2025-10-01 HackerOne reports 560+ valid reports from autonomous agents
2025-10-02 Benjamin Jones’s AI-in-R&D model issued as NBER w34312
2025-10-05 GDPval: 47.6% wins-or-ties against experts on a 220-task subset
2025-10-23 A human beats AlphaEvolve’s kissing-number bound, five months on
2025-11-03 AlphaEvolve on 67 mathematical problems; analytic number theory the recorded failure
2025-11-08 SWE-fficiency, arXiv
2025-11-12 AlphaProof published in Nature; Ringer’s hands-on assessment appears
2025-11-13 Brynjolfsson, Chandar, and Chen: entry-level employment down 16%
2025-11-20 Early science acceleration experiments with GPT-5, arXiv
2025-11-22 Tao on problem selection, Mathstodon
2025-11-28 ThetaEvolve: an 8B open model passes two AlphaEvolve bounds
2025-11-30 Tao on the long tail of unsolved problems, Mathstodon
2025-12-08 The Equational Theories Project settles 22,028,942 implications
2025-12-10 ARTEMIS pentest comparison, arXiv
2026-01-07 Tao’s caveat on Erdős #728, five days before the writeup
2026-01-12 Erdős #728 Lean-proof writeup, arXiv
2026-01-16 Locus/Intology modded-nanogpt record, 1.765 min
2026-01-21 curl ends its bug bounty after AI submission flood
2026-01-22 TTT-Discover, arXiv
2026-01-29 METR Time Horizon 1.1: recent doubling time falls to 88.6 days
2026-02 Opus 4.6, end of the AISI cyber-range series (1.7 → 9.8 steps)
2026-02-02 Aster modded-nanogpt record, 1.528 min
2026-02-10 Station modded-nanogpt record, 1.496 min
2026-02-13 NBER w34851: AI closes three quarters of the education productivity gap
2026-02-18 Simple baselines match code evolution on nine AlphaEvolve problems
2026-03-01 Jagged-frontier experiment published in Organization Science
2026-03-06 Karpathy’s autoresearch repository created
2026-03-11 UK AISI multi-step cyber attack scenarios, arXiv
2026-03-16 Gemini and developer experience: experience not substitutable for code security
2026-03-16 HorizonMath: 100+ unsolved problems, models near 0%
2026-03-20 Tao’s Dwarkesh interview: “jumping machines”
2026-03-29 Tao’s blog post on AI as a complementary style
2026-04 Anthropic previews Mythos; thousands of claimed vulnerabilities
2026-04-02 Lyptus offensive-cyber time horizons
2026-04-02 Bazzichi, Riccaboni, and Castellacci on recombinant innovation
2026-04-07 AISLE: no stable best model across cyber tasks
2026-04-09 Vidoc reproduces Mythos-class detection with public models
2026-04-14 Breunig on UK AISI Mythos runs: no diminishing returns at 100M tokens
2026-04-30 Davidson, Halperin, Houlden, and Korinek, NBER w35155
2026-05-07 AlphaEvolve one-year update: ~0.7% of fleet compute
2026-05-13 Axios follow-up on the Palo Alto Mythos deployment
2026-05-20 Unit-distance conjecture disproved; verification posted the same day
2026-05-21 Current modded-nanogpt record, 1.320 min (~34× the baseline)
2026-05-21 AlphaProof Nexus: 9 of 353 Erdős problems, 44 of 492 OEIS conjectures
2026-05-26 AppSec job-posting analysis: AI mentions rise 2.1% to 7.2% in six months
2026-05-27 GPT-5.5 saturates the offensive-cyber horizon task set
2026-04-15 NIST: CVE submissions up 263% since 2020, enrichment cannot keep pace
2026-05-28 Williams contextualizes the unit-distance disproof
2026-06-12 FrontierMath v2 released after errors in 42% of problems
2026-06-30 Erdős-problems AI wiki freezes
2026-07 DeepMind’s validation-bottleneck essay
2026-07-01 Performance-benchmark reliability audit, arXiv
2026-07-08 PERFOPT-Bench relay pilot, arXiv
2026-07-10 Kenyan entrepreneur RCT pre-published in Management Science
2026-07-22 SANS 2026 workforce findings reported: 16% cut headcount, entry level hardest hit
2026-07-26 This log last checked; the syntheses assembled, including the exponent series and the AlphaEvolve record sample

References

Acemoglu, Daron. 2024. “The Simple Macroeconomics of AI.” National Bureau of Economic Research. https://economics.mit.edu/sites/default/files/2024-04/The%20Simple%20Macroeconomics%20of%20AI.pdf.
Aghion, Philippe, and Simon Bunel. 2024. “AI and Growth: Where Do We Stand.” https://www.frbsf.org/wp-content/uploads/AI-and-Growth-Aghion-Bunel.pdf.
Aghion, Philippe, Benjamin F. Jones, and Charles I. Jones. 2019. “Artificial Intelligence and Economic Growth.” In The Economics of Artificial Intelligence: An Agenda, edited by Ajay Agrawal, Joshua Gans, and Avi Goldfarb, 237–90. Chicago: University of Chicago Press. https://www.degruyterbrill.com/document/doi/10.7208/9780226613475-011/html?lang=en.
Besiroglu, Tamay, Nicholas Emery-Xu, and Neil Thompson. 2024. “Economic Impacts of AI-Augmented r&d.” Research Policy 53 (7): 105037. https://doi.org/10.1016/j.respol.2024.105037.
Brynjolfsson, Erik, Bharat Chandar, and Daniel Chen. 2025. “Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence.” Stanford Digital Economy Lab. https://digitaleconomy.stanford.edu/wp-content/uploads/2025/08/Canaries_BrynjolfssonChandarChen.pdf.
Folkerts, Linus, Will Payne, Simon Inman, Philippos Giavridis, Joe Skinner, Sam Deverett, James Aung, et al. 2026. “Measuring AI Agents’ Progress on Multi-Step Cyber Attack Scenarios.” https://arxiv.org/abs/2603.11214.
Ide, Enrique, and Eduard Talamas. 2024. “Artificial Intelligence in the Knowledge Economy.” https://doi.org/10.1086/737233.
Jones, Benjamin F. 2025. “Artificial Intelligence in Research and Development.” NBER Working Paper 34312. National Bureau of Economic Research. https://doi.org/10.3386/w34312.
Jordan, Keller, and contributors. 2026. “Modded-Nanogpt.” https://github.com/KellerJordan/modded-nanogpt.
Korinek, Anton, and Donghyun Suh. 2024. “Scenarios for the Transition to AGI.” National Bureau of Economic Research. https://arxiv.org/pdf/2403.12107.pdf.
Ma, Jeffrey Jian, Milad Hashemi, Amir Yazdanbakhsh, et al. 2025. SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?” https://arxiv.org/pdf/2511.06090.pdf.
Nathani, Deepak, Lovish Madaan, Nicholas Roberts, et al. 2025. MLGym: A New Framework and Benchmark for Advancing AI Research Agents.” https://arxiv.org/pdf/2502.14499.pdf.
Novikov, Alexander, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, et al. 2025. “AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery.” arXiv Preprint arXiv:2506.13131. https://doi.org/10.48550/arXiv.2506.13131.
Noy, Shakked, and Whitney Zhang. 2023. “Experimental Evidence on the Productivity Effects of Generative AI.” Science 381 (6654): 187–92. https://doi.org/10.1126/science.adh2586.
OpenAI. 2026. “An OpenAI Model Has Disproved a Central Conjecture in Discrete Geometry.” May 20, 2026. https://openai.com/index/model-disproves-discrete-geometry-conjecture/.
Press, Ori, Brandon Amos, Haoyu Zhao, et al. 2025. AlgoTune: Can Language Models Speed up General-Purpose Numerical Programs?” https://arxiv.org/pdf/2507.15887.pdf.
Shetty, Manish, Naman Jain, Jinjian Liu, Vijay Kethanaboyina, Koushik Sen, and Ion Stoica. 2025. “GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents.” https://arxiv.org/pdf/2505.23671.pdf.
Wijk, Hjalmar, Tao Lin, Joel Becker, Sami Jawhar, Neev Parikh, Thomas Broadley, Lawrence Chan, et al. 2025. “RE-Bench: Evaluating Frontier AI r&d Capabilities of Language Model Agents Against Human Experts.” https://arxiv.org/abs/2411.15114.
Yuksekgonul, Mert, Daniel Koceja, Xinhao Li, Federico Bianchi, Jed McCaleb, Xiaolong Wang, Jan Kautz, et al. 2026. “Learning to Discover at Test Time.” arXiv Preprint arXiv:2601.16175. https://test-time-training.github.io/discover.pdf.