- How much have LLMs accelerated discovery?
-
My very tentative conclusions:
- Discovery of vulnerabilities has accelerated sharply.
- Discovery in mathematics has likely accelerated, but it’s hard to benchmark.
- Discovery of optimizations has not shown a clear acceleration (surprising to me!).
- Discovery in other sciences are very hard to benchmark.
- Most LLM-assisted discoveries did not require domain expertise (i.e. semi-autonomous).
- Most LLM-assisted discoveries have a low average cost per discovery.
- We are just looking for slope changes.
- AI-assisted discoveries are announced every afternoon, but the significance is hard to assess. Here will just look for slope changes where we have historical data on a regular pace of discovery. For concreteness I treat a slope change around January 2026 to be primarily due to LLMs.
- Discovery of vulnerabilities has accelerated sharply.
- Across curl, OpenSSL, and Firefox, the reporting of vulnerabilities accelerated, and they had high shares that were explicitly attributed to AI (42%, 66%, 12%).
- Discovery in mathematics has likely accelerated, but hard to benchmark.
- Several prominent AI-attributed discoveries, but no measured aggregate acceleration yet. The century-scale exponent series are flat through the agent era. The Erdős catalogue shows recent status churn, but it provides only about eleven months of comparable snapshots and the database doesn’t have reliable solution dates so it is hard to establish an historical trend.
- Discovery of optimizations has not show a marked acceleration.
- We have good historical time-series for algorithmic efficiency or quality across many domains, and there does not seem to be a significant speedup in 2026, despite LLM making contributions to SAT solving, small-matrix multiplication, ML speedruns, and GPU kernels.
- Discovery is not expertise-loaded.
- Many of the LLM-assisted discoveries below appear to be fairly autonomous, i.e. most were produced by people who were not domain experts in the problem. This is different from normal progress in these domains, where discoveries are typically made only by people who have very deep expertise.
- LLM-assisted discoveries are cheap.
- The average cost/discovery appears to be lower for autonomous LLM-produced discoveries, compared to human-produced discoveries. The average cost can be hard to interpret, but a simple assumption is that people are spending on LLM test-time compute until the marginal returns to expenditure are roughly equalized between humans and LLMs. In this case a lower average cost implies a lower elasticity (the elasticity is the ratio of marginal to average returns).
- Why is this important to know?
- It’s particularly important for anticipating when LLMs will be making significant contributions to AI research.
Discovery in Vulnerabilities
Notable LLM Contributions
| Date | System / source | Real-world findings | Setting / evidence type | Expenditure |
|---|---|---|---|---|
| Nov 2024–Aug 2025 | Google Big Sleep | Found CVE-2025-6965 in production SQLite before exploitation; separately disclosed 20 previously unknown bugs in widely used OSS including FFmpeg and ImageMagick. An earlier SQLite find was confined to a development branch | Agent found and reproduced the bugs autonomously; humans handled final review and disclosure; vendor self-report | Unknown |
| Jun 2025 | XBOW on HackerOne | About 132 bug-bounty reports were confirmed and resolved on live customer web applications, from about 1{,}060 submissions; ~208 were duplicates and ~209 merely informative | Autonomous pentester; vendor figures, not independently audited here; confirmed reports are not all established zero-days | Unknown |
| Aug 2025 | DARPA AIxCC finalists | Found 18 genuine, non-synthetic zero-days in competition OSS codebases (6 C, 12 Java) and supplied patches for 11 | Seven systems ran without human intervention during the competition; real OSS in a competition environment, not a production deployment | Unknown for the genuine zero-days |
| Dec 2025 | ARTEMIS | Discovered 9 valid vulnerabilities on a live university network of about 8{,}000 hosts; one was found only by the two agent variants | Purpose-built academic scaffold in a controlled head-to-head study | Some variants cost about $18/hour; total finding cost unknown |
| Mar 2026 | CyberGym | During runs aimed at reproducing described vulnerabilities, agents incidentally uncovered 34 zero-days and 18 historically incomplete patches in real OSS | Academic reproduction benchmark; new findings were a by-product rather than the primary task | Unknown |
| Apr 2026 | Anthropic Mythos | Vendor reports “thousands” of previously unknown vulnerabilities; public examples include a 27-year-old OpenBSD DoS and FFmpeg flaws fixed in version 8.1 | Large-scale scaffolded search; aggregate is an unaudited vendor claim | OpenBSD example: <$20{,}000 across about 1{,}000 runs; successful run <$50 |
These rows are not additive. The same systems, teams, and disclosures recur in project records and vendor totals.
AI-credited disclosures on fixed codebases
| Project record | 2026 cutoff | AI-credited / all disclosures | Important qualification |
|---|---|---|---|
| curl | Jun 24 | 15 / 36 (42%) | AI-credited issues were 80% Low and none High or Critical; the share is a floor based on explicit finder credits |
| OpenSSL | Jun 9 | 25 / 38 (66%) | Highest AI-credited share here, but the same researchers appear in other project and system rows |
| Firefox | late Jul | 137 / 1{,}139 CVEs (12%) | 121 of 137 credit one seven-person Anthropic team; advisory counts are disclosures, not clean discovery counts |
Discovery in Math
Solving pre-specified problems
| List | Problems | Resolved | Resolved with LLMs |
|---|---|---|---|
| Erdős problems (community catalogue; ongoing) | 1,217 catalogued | 565 marked solved, but status dates are not solution dates | ~13 full AI-standalone resolutions (wiki freeze 2026-06-30) ~50 assisted resolutions (Tao’s separate informal count, not a rate) |
| Hilbert problems (1900) | 23 | ~12 consensus; last major dated piece Hales 1998 | 0 |
| Smale problems (1998) | 18 | 5 | 1, Jacobian conjecture counterexample, Jul 2026; formally checked, peer review pending |
| Millennium problems (2000) | 7 | 1 (Poincaré, Perelman 2003) | 0 |
| TOPP (2001–) | 78 | 17 marked solved/settled/closed | 0 |
Tightening bounds
| Bound | How many | Earliest → latest (example) | Steps | AI? | Expertise |
|---|---|---|---|---|---|
| Analytic number theory exponents (ANTEDB) | 3 families (\(\mu\), \(A\), \(\beta\)); 58 slices (\(20{+}19{+}19\)) | Lindelöf \(\mu(1/2)\): 1920 at \(5/28\approx0.179\) → 2017 at \(13/84\approx0.155\) | 15 on Lindelöf | no LLM step; ANTEDB’s non-LLM linear-programming collation did improve some bounds at launch | — |
| Sphere-packing density | 1 lower-bound ladder (form changes); Astra claims on upper side | Minkowski–Hlawka 1905 → Klartag 2025 (\(c\,n^{2}2^{-n}\)) | 8 (lower) | claimed on upper only (Astra Aug 2026; Lean; peer review pending) | vendor / AI lab |
| Kissing number (dim. 11) | 1 dimension among the kissing family | 1971 at 566 → 2026 at 604 | 5 | yes: AlphaEvolve 593, then the same agent route 594 and 604 | mixed (specialist lab and mathematicians; later collective agents) |
| Sums-and-differences / autoconvolution | 2 tracked ladders (of 3+5 problems in those AlphaEvolve topic groups) | \(C_{6.44}\): 2007 \(\approx1.079\) → 2025 \(\approx1.173\); \(C_{6.3}\): 2010 \(0.889\) → 2025 \(0.961\) | 8; 4 | yes; humans retook \(C_{6.44}\), while AlphaEvolve’s larger \(C_{6.3}\) step followed Boyer–Li’s result | specialist (AlphaEvolve + math coauthors) |
| Matrix-multiplication exponent \(\omega\) | 1 asymptotic exponent | Four small human improvements in the past ~14 years | 4 recent | no LLM step on the asymptotic ladder | — |
| Difference-basis / hexagon / min-overlap type | finite geometric/packing constants in the AlphaEvolve inventory (65 numbered; ~50 distinct) | e.g. difference-basis: Golay 1972 at \(\approx2.657\) → AlphaEvolve 2025 at \(2.639\); min-overlap 1955→2025 | 4; 5 | yes on a minority; the difference-basis step required Singer-code hints | specialist (AlphaEvolve + math coauthors) |
Notable LLM Contributions
| Date | Result | What happened | Domain expertise / prompt generality | Expenditure |
|---|---|---|---|---|
| Dec 2023 | FunSearch (cap set) | Improved asymptotic cap-set lower bound (claimed largest in ~20 years); LLM evolves programs scored by an automated evaluator | High scaffolding: problem recast as program search with a custom evaluator | Few dozen iterations over a few days; $ unknown |
| May 2025 | AlphaEvolve (e.g. kissing # in dim. 11) | Raised dim.-11 kissing lower bound 592→593; later paper: improved a minority of ~50–67 open problems (mostly small nudges) | High: needs a computable construction objective; specialist coauthors; expert hints on some targets | Unknown |
| May 2026 | OpenAI unit-distance disproof | Disproved Erdős’s 1946 unit-distance growth conjecture (\(n^{1+\delta}\) construction); independently digested by nine mathematicians | Low (vendor claim): general-purpose model, not specialized for math, not scaffolded for proof search, not targeted at this problem | Unknown |
| May 2026 | AlphaProof Nexus | Formal verify-and-retry agents resolved 9 of a vendor-defined set of 353 open Erdős problems (~2.5%); also 44/492 OEIS conjectures | Medium: scaffolded Lean search on a formalizable corpus, not unaided research agenda-setting | ~few hundred $ each (vendor) |
| Jul 2026 | Alpöge–Claude Fable (Smale 16) | Counterexample to the Jacobian conjecture in dim. \(\ge3\); independently formalized in Lean and Isabelle; peer review pending | High human expertise: domain specialist directing a frontier chat model | Unknown |
| Aug 2026 | OpenAI Astra | Vendor package of claimed advances with Lean certificates, including packing/coding upper bounds and Erdős #146/#180/#183; peer review pending | AI-lab hybrid; the source log does not establish a comparable human-input denominator | Unknown |
Discovery in Algorithms
Notable LLM Contributions
| Date | Result | What happened | Domain expertise / prompt generality | Expenditure |
|---|---|---|---|---|
| Aug 2024 | The AI Scientist (Sakana) | End-to-end loop claiming full ML papers from ideas; quality judged mainly by the authors’ own automated reviewer | Medium–high scaffolding: templated experiment/code pipelines (v1); not free-form research | <$15 per paper (vendor) |
| May 2025 | AlphaEvolve | Improved ~20% and matched ~75% of DeepMind’s 50-plus math-and-algorithms targets; found a 48-multiply 4×4 complex matmul; a 23% kernel speedup translated to ~1% of Gemini training time, while production heuristics saved ~0.7% of fleet compute | High: needs a searchable program and computable objective; specialist lab system | Unknown (production value is press/vendor, not a finding cost) |
| Aug 2025 | AE-Kissat-MAB (SAT Competition) | LLM-evolved solver won by 6 of 400 instances (~2%) over the human-written runner-up | Human-steered evolutionary rounds; the changing annual task set prevents placing the win on the fixed-hardware long-run curve | Unknown |
| Oct 2025 | CIFAR-10 speedrun (Hiverge) | First acknowledged AI-set record: 94% in 1.99 s on one A100 (~23% step); a later 1.828 s claim remains unacknowledged and has specification-gaming caveats | High scaffolding: discovery engine aimed at a fixed speedrun metric; series otherwise flattening | Unknown |
| Jan 2026 | TTT-Discover | A test-time-RL harness using open gpt-oss-120b found TriMul GPU kernels 15–51% faster than the best human kernels, depending on hardware | The frontier step came from scaffolding and test-time training, not a new base-model generation | ~few hundred $ of test-time compute (authors; accounting incomplete) |
| Mar 2026 | Karpathy autoresearch | In ~2 days it worked through ~700 nanochat changes, about 20 of which improved validation loss; Karpathy later tested, transferred, and stacked them and measured an ~11% reduction in “time to GPT-2” | Autonomous mechanical search during the run; specialist evaluation and integration afterward | Unknown (single-GPU ~2-day run) |
Notes
- We’re mainly looking at LLMs, not computers or AI in general.
- Most science can’t be judged with a single number.
-
The work of most scientists can’t be judged with a single scalar. Physicists, chemists, biologists, typically are judged by their peers on qualitative grounds.
When we do have a single number (either a set of problems, an upper or lower bound, or a measure of efficiency), then these are generally based on problems with cheap validation.
Unit cost over time for a selection of the 66 technologies in the Santa Fe Performance Curve Database. Log vertical axis, so a straight line is a constant exponential rate. Each series is priced in its own unit, so only slopes are comparable, never levels. These series end in 2013 and are a pre-AI baseline. - Paper and code production increased sharply.
- Monthly arXiv submissions rose from 17{,}271 in November 2022 to 32{,}040 in June 2026, about 85%. GitHub pushes roughly doubled from late 2024 to early 2026. Neither series measures quality: output volume can rise substantially without comparable growth in genuine discovery.
- Assessing the significance of individual discoveries is hard.
- A new record can be a deep conceptual advance or a tiny movement on a neglected target. Without the size of the step, the prior trend, and a denominator of comparable attempts or human results, a list of successes cannot establish aggregate acceleration.
- Measuring discovery ability is hard.
- In principle we could measure an LLM’s discovery ability offline by giving it a problem and asking for a new result. In practice success and failure are difficult to interpret because of problem selection and significance, contamination, inference-time scaling, and the division of work between models, scaffolds, and human experts. Time horizons and cost per result are useful denominators, but neither by itself measures the value of a discovery.



