Draft

LLMs’ Contribution to Discovery

Author
Affiliation

Tom Cunningham

METR

Published

August 8, 2026

How much have LLMs accelerated discovery?

My very tentative conclusions:

  1. Discovery of vulnerabilities has accelerated sharply.
  2. Discovery in mathematics has likely accelerated, but it’s hard to benchmark.
  3. Discovery of optimizations has not shown a clear acceleration (surprising to me!).
  4. Discovery in other sciences are very hard to benchmark.
  5. Most LLM-assisted discoveries did not require domain expertise (i.e. semi-autonomous).
  6. Most LLM-assisted discoveries have a low average cost per discovery.
We are just looking for slope changes.
AI-assisted discoveries are announced every afternoon, but the significance is hard to assess. Here will just look for slope changes where we have historical data on a regular pace of discovery. For concreteness I treat a slope change around January 2026 to be primarily due to LLMs.
Discovery of vulnerabilities has accelerated sharply.
Across curl, OpenSSL, and Firefox, the reporting of vulnerabilities accelerated, and they had high shares that were explicitly attributed to AI (42%, 66%, 12%).
Discovery in mathematics has likely accelerated, but hard to benchmark.
Several prominent AI-attributed discoveries, but no measured aggregate acceleration yet. The century-scale exponent series are flat through the agent era. The Erdős catalogue shows recent status churn, but it provides only about eleven months of comparable snapshots and the database doesn’t have reliable solution dates so it is hard to establish an historical trend.
Discovery of optimizations has not show a marked acceleration.
We have good historical time-series for algorithmic efficiency or quality across many domains, and there does not seem to be a significant speedup in 2026, despite LLM making contributions to SAT solving, small-matrix multiplication, ML speedruns, and GPU kernels.
Discovery is not expertise-loaded.
Many of the LLM-assisted discoveries below appear to be fairly autonomous, i.e. most were produced by people who were not domain experts in the problem. This is different from normal progress in these domains, where discoveries are typically made only by people who have very deep expertise.
LLM-assisted discoveries are cheap.
The average cost/discovery appears to be lower for autonomous LLM-produced discoveries, compared to human-produced discoveries. The average cost can be hard to interpret, but a simple assumption is that people are spending on LLM test-time compute until the marginal returns to expenditure are roughly equalized between humans and LLMs. In this case a lower average cost implies a lower elasticity (the elasticity is the ratio of marginal to average returns).
Why is this important to know?
It’s particularly important for anticipating when LLMs will be making significant contributions to AI research.

Discovery in Vulnerabilities

Six bar charts in two rows, all with annual counts on linear axes and January 1, 2026 onward shaded. Top row: curl, rising from about ten a year to 36 in a part-year 2026 with 15 credited to AI; OpenSSL, a mid-2010s peak of 35 then single digits, then 38 in part-2026 of which 25 are credited to AI; Firefox, rising from 119 in 2016 to 640 in 2025 and 1,139 in part-2026, with a thin amber fuzzer band and a red AI band of 137. Bottom row: OSS-Fuzz falling from 1,041 in 2020 to 244 in 2025; all CVEs disclosed in the US national database rising from 6,500 in 2016 to about 50,000 in 2025; and additions to the government exploited-vulnerabilities catalogue, roughly flat between 170 and 550 a year since the catalogue began in late 2021.

Six vulnerability-discovery series in one format: three fixed codebases that credit their finders, an automated fuzzing programme, and all software split into found and exploited.

Notable LLM Contributions

Date System / source Real-world findings Setting / evidence type Expenditure
Nov 2024–Aug 2025 Google Big Sleep Found CVE-2025-6965 in production SQLite before exploitation; separately disclosed 20 previously unknown bugs in widely used OSS including FFmpeg and ImageMagick. An earlier SQLite find was confined to a development branch Agent found and reproduced the bugs autonomously; humans handled final review and disclosure; vendor self-report Unknown
Jun 2025 XBOW on HackerOne About 132 bug-bounty reports were confirmed and resolved on live customer web applications, from about 1{,}060 submissions; ~208 were duplicates and ~209 merely informative Autonomous pentester; vendor figures, not independently audited here; confirmed reports are not all established zero-days Unknown
Aug 2025 DARPA AIxCC finalists Found 18 genuine, non-synthetic zero-days in competition OSS codebases (6 C, 12 Java) and supplied patches for 11 Seven systems ran without human intervention during the competition; real OSS in a competition environment, not a production deployment Unknown for the genuine zero-days
Dec 2025 ARTEMIS Discovered 9 valid vulnerabilities on a live university network of about 8{,}000 hosts; one was found only by the two agent variants Purpose-built academic scaffold in a controlled head-to-head study Some variants cost about $18/hour; total finding cost unknown
Mar 2026 CyberGym During runs aimed at reproducing described vulnerabilities, agents incidentally uncovered 34 zero-days and 18 historically incomplete patches in real OSS Academic reproduction benchmark; new findings were a by-product rather than the primary task Unknown
Apr 2026 Anthropic Mythos Vendor reports “thousands” of previously unknown vulnerabilities; public examples include a 27-year-old OpenBSD DoS and FFmpeg flaws fixed in version 8.1 Large-scale scaffolded search; aggregate is an unaudited vendor claim OpenBSD example: <$20{,}000 across about 1{,}000 runs; successful run <$50

These rows are not additive. The same systems, teams, and disclosures recur in project records and vendor totals.

AI-credited disclosures on fixed codebases

Project record 2026 cutoff AI-credited / all disclosures Important qualification
curl Jun 24 15 / 36 (42%) AI-credited issues were 80% Low and none High or Critical; the share is a floor based on explicit finder credits
OpenSSL Jun 9 25 / 38 (66%) Highest AI-credited share here, but the same researchers appear in other project and system rows
Firefox late Jul 137 / 1{,}139 CVEs (12%) 121 of 137 credit one seven-person Anthropic team; advisory counts are disclosures, not clean discovery counts

Discovery in Math

Eight panels in two rows, each a step function with years on the x-axis and January 1, 2026 onward shaded. Panel 1: thirteen Erdős-database status snapshots spanning about eleven months, with problems catalogued rising to 1217, recorded solved statuses rising to 565, and Lean-formalized statements rising from 148 to 605. The window from April 2026, when the catalogue count stopped changing, is separately shaded; a single red point at 13 marks the AI-standalone full resolutions in the project's June 2026 wiki freeze. Status dates are not solution dates, and the two stocks are not an AI-versus-human flow comparison. Panel 2: the sums-and-differences lower bound rising from 1.079 to 1.173, four human steps in 2007, then in 2025 two AlphaEvolve steps and two human steps above them. Panel 3: the autoconvolution lower bound, one human step in 2010 and then three leapfrogging steps in 2025, AlphaEvolve, a human gradient method, then AlphaEvolve again. Panel 4: the kissing number in dimension 11, human steps to 592 in 2022, AlphaEvolve to 593 in 2025, then collective agents to 604 in 2026. Panel 5: the Lindelöf exponent descending from 0.179 to 0.155 over fifteen steps between 1920 and 2017, flat through the shaded 2026 period. Panel 6: the mu slice at three fifths, seventeen small human steps from 1920 to 2023. Panel 7: the zero-density exponent at three quarters, flat from Ingham in 1940 until a human improvement in 2024. Panel 8: the beta exponent at one tenth, five steps between 1993 and 2017 and flat since.

The math evidence on the two overview questions: AI-attributed stocks, targeted rates, and record steps where comparison is possible (top row), and the long-run exponent records that show no change of slope (bottom row).

Solving pre-specified problems

List Problems Resolved Resolved with LLMs
Erdős problems (community catalogue; ongoing) 1,217 catalogued 565 marked solved, but status dates are not solution dates ~13 full AI-standalone resolutions (wiki freeze 2026-06-30)
~50 assisted resolutions (Tao’s separate informal count, not a rate)
Hilbert problems (1900) 23 ~12 consensus; last major dated piece Hales 1998 0
Smale problems (1998) 18 5 1, Jacobian conjecture counterexample, Jul 2026; formally checked, peer review pending
Millennium problems (2000) 7 1 (Poincaré, Perelman 2003) 0
TOPP (2001–) 78 17 marked solved/settled/closed 0

Tightening bounds

Bound How many Earliest → latest (example) Steps AI? Expertise
Analytic number theory exponents (ANTEDB) 3 families (\(\mu\), \(A\), \(\beta\)); 58 slices (\(20{+}19{+}19\)) Lindelöf \(\mu(1/2)\): 1920 at \(5/28\approx0.179\) → 2017 at \(13/84\approx0.155\) 15 on Lindelöf no LLM step; ANTEDB’s non-LLM linear-programming collation did improve some bounds at launch
Sphere-packing density 1 lower-bound ladder (form changes); Astra claims on upper side Minkowski–Hlawka 1905 → Klartag 2025 (\(c\,n^{2}2^{-n}\)) 8 (lower) claimed on upper only (Astra Aug 2026; Lean; peer review pending) vendor / AI lab
Kissing number (dim. 11) 1 dimension among the kissing family 1971 at 566 → 2026 at 604 5 yes: AlphaEvolve 593, then the same agent route 594 and 604 mixed (specialist lab and mathematicians; later collective agents)
Sums-and-differences / autoconvolution 2 tracked ladders (of 3+5 problems in those AlphaEvolve topic groups) \(C_{6.44}\): 2007 \(\approx1.079\) → 2025 \(\approx1.173\); \(C_{6.3}\): 2010 \(0.889\) → 2025 \(0.961\) 8; 4 yes; humans retook \(C_{6.44}\), while AlphaEvolve’s larger \(C_{6.3}\) step followed Boyer–Li’s result specialist (AlphaEvolve + math coauthors)
Matrix-multiplication exponent \(\omega\) 1 asymptotic exponent Four small human improvements in the past ~14 years 4 recent no LLM step on the asymptotic ladder
Difference-basis / hexagon / min-overlap type finite geometric/packing constants in the AlphaEvolve inventory (65 numbered; ~50 distinct) e.g. difference-basis: Golay 1972 at \(\approx2.657\) → AlphaEvolve 2025 at \(2.639\); min-overlap 1955→2025 4; 5 yes on a minority; the difference-basis step required Singer-code hints specialist (AlphaEvolve + math coauthors)

Notable LLM Contributions

Date Result What happened Domain expertise / prompt generality Expenditure
Dec 2023 FunSearch (cap set) Improved asymptotic cap-set lower bound (claimed largest in ~20 years); LLM evolves programs scored by an automated evaluator High scaffolding: problem recast as program search with a custom evaluator Few dozen iterations over a few days; $ unknown
May 2025 AlphaEvolve (e.g. kissing # in dim. 11) Raised dim.-11 kissing lower bound 592→593; later paper: improved a minority of ~50–67 open problems (mostly small nudges) High: needs a computable construction objective; specialist coauthors; expert hints on some targets Unknown
May 2026 OpenAI unit-distance disproof Disproved Erdős’s 1946 unit-distance growth conjecture (\(n^{1+\delta}\) construction); independently digested by nine mathematicians Low (vendor claim): general-purpose model, not specialized for math, not scaffolded for proof search, not targeted at this problem Unknown
May 2026 AlphaProof Nexus Formal verify-and-retry agents resolved 9 of a vendor-defined set of 353 open Erdős problems (~2.5%); also 44/492 OEIS conjectures Medium: scaffolded Lean search on a formalizable corpus, not unaided research agenda-setting ~few hundred $ each (vendor)
Jul 2026 Alpöge–Claude Fable (Smale 16) Counterexample to the Jacobian conjecture in dim. \(\ge3\); independently formalized in Lean and Isabelle; peer review pending High human expertise: domain specialist directing a frontier chat model Unknown
Aug 2026 OpenAI Astra Vendor package of claimed advances with Lean certificates, including packing/coding upper bounds and Erdős #146/#180/#183; peer review pending AI-lab hybrid; the source log does not establish a comparable human-input denominator Unknown

Discovery in Algorithms

Eight panels in a two-by-four grid, each a step-function time series with January 1, 2026 onward shaded. Panel 1: modded-nanogpt training minutes on a log axis falling from 45 in May 2024 to 1.27 in May 2026 over 86 records, with four red AI-set records in the flat tail. Panel 2: CIFAR-10 speedrun seconds to 94 percent on a log axis falling from 18.1 in December 2022 to 1.99 in October 2025, the last filled record red and a July 2026 open red dot for an unacknowledged 1.828-second claim. Panel 3: Stockfish Elo against Stockfish 15 rising from minus 538 in 2013 to plus 137 in July 2026, with an NNUE jump in 2020 and an open red circle marking the first LLM-credited commit. Panel 4: Hutter Prize enwik9 total size falling from 116.7 to 110.8 megabytes over human records through 2024, a pending 2026 entry as an open dot, and a dashed line for the uncapped leaderboard flat since October 2023. Panel 5: the enwik8 prize, four records over 2006 to 2017, ending before the highlighted period. Panel 6: Gurobi MILP cumulative vendor speedup stepping from 1 to about 1.5 over 2022 to 2025, drawn grey. Panel 7: the matrix multiplication exponent staircase from 2.8074 in 1969 to 2.371339 in 2024, every step human. A figure-level key identifies human, AI, vendor-run, and pending records. Panel 8 states the flow and slope verdicts.

Seven record series in one format: years on the x-axis, the standing best on the y-axis, every known step drawn, January 1, 2026 onward shaded, AI-set records in red. The top row holds the series AI has entered; the bottom row the series it has not.

Notable LLM Contributions

Date Result What happened Domain expertise / prompt generality Expenditure
Aug 2024 The AI Scientist (Sakana) End-to-end loop claiming full ML papers from ideas; quality judged mainly by the authors’ own automated reviewer Medium–high scaffolding: templated experiment/code pipelines (v1); not free-form research <$15 per paper (vendor)
May 2025 AlphaEvolve Improved ~20% and matched ~75% of DeepMind’s 50-plus math-and-algorithms targets; found a 48-multiply 4×4 complex matmul; a 23% kernel speedup translated to ~1% of Gemini training time, while production heuristics saved ~0.7% of fleet compute High: needs a searchable program and computable objective; specialist lab system Unknown (production value is press/vendor, not a finding cost)
Aug 2025 AE-Kissat-MAB (SAT Competition) LLM-evolved solver won by 6 of 400 instances (~2%) over the human-written runner-up Human-steered evolutionary rounds; the changing annual task set prevents placing the win on the fixed-hardware long-run curve Unknown
Oct 2025 CIFAR-10 speedrun (Hiverge) First acknowledged AI-set record: 94% in 1.99 s on one A100 (~23% step); a later 1.828 s claim remains unacknowledged and has specification-gaming caveats High scaffolding: discovery engine aimed at a fixed speedrun metric; series otherwise flattening Unknown
Jan 2026 TTT-Discover A test-time-RL harness using open gpt-oss-120b found TriMul GPU kernels 15–51% faster than the best human kernels, depending on hardware The frontier step came from scaffolding and test-time training, not a new base-model generation ~few hundred $ of test-time compute (authors; accounting incomplete)
Mar 2026 Karpathy autoresearch In ~2 days it worked through ~700 nanochat changes, about 20 of which improved validation loss; Karpathy later tested, transferred, and stacked them and measured an ~11% reduction in “time to GPT-2” Autonomous mechanical search during the run; specialist evaluation and integration afterward Unknown (single-GPU ~2-day run)

Notes

We’re mainly looking at LLMs, not computers or AI in general.
Most science can’t be judged with a single number.

The work of most scientists can’t be judged with a single scalar. Physicists, chemists, biologists, typically are judged by their peers on qualitative grounds.

When we do have a single number (either a set of problems, an upper or lower bound, or a measure of efficiency), then these are generally based on problems with cheap validation.

Unit cost over time for a selection of the 66 technologies in the Santa Fe Performance Curve Database. Log vertical axis, so a straight line is a constant exponential rate. Each series is priced in its own unit, so only slopes are comparable, never levels. These series end in 2013 and are a pre-AI baseline.

Unit cost over time for a selection of the 66 technologies in the Santa Fe Performance Curve Database. Log vertical axis, so a straight line is a constant exponential rate. Each series is priced in its own unit, so only slopes are comparable, never levels. These series end in 2013 and are a pre-AI baseline.
Paper and code production increased sharply.
Monthly arXiv submissions rose from 17{,}271 in November 2022 to 32{,}040 in June 2026, about 85%. GitHub pushes roughly doubled from late 2024 to early 2026. Neither series measures quality: output volume can rise substantially without comparable growth in genuine discovery.
Assessing the significance of individual discoveries is hard.
A new record can be a deep conceptual advance or a tiny movement on a neglected target. Without the size of the step, the prior trend, and a denominator of comparable attempts or human results, a list of successes cannot establish aggregate acceleration.
Measuring discovery ability is hard.
In principle we could measure an LLM’s discovery ability offline by giving it a problem and asking for a new result. In practice success and failure are difficult to interpret because of problem selection and significance, contamination, inference-time scaling, and the division of work between models, scaffolds, and human experts. Time horizons and cost per result are useful denominators, but neither by itself measures the value of a discovery.