Draft

Apple-Picking Offcuts

2026-06-03 | econbench

2026-06-01 | expenditure

  • Humans find 5 bugs/month (unstated expenditure)
  • Mythos found 24 bugs at $1M expenditure, $40,000/bug

https://x.com/tim_hua_/status/2061498936813322663?s=46

“Mythos at Palo Alto Networks”found more than two dozen critical vulnerabilities in around three weeks, roughly five times what the company would typically find using existing tools … But the company “burned through more than $1 million worth of tokens using Mythos””

https://www.theinformation.com/articles/anthropics-mythos-security-powerhouse-budget-buster

More cost/economics details (relevant to the expenditure-horizon picture):

  • Per-find cost is highly variable. In Anthropic’s own testing, ~1,000 scaffold runs on OpenBSD cost under $20,000 total and turned up several dozen findings; the single run that found the headline critical bug cost under $50 — but which run succeeds isn’t knowable in advance, so the expected cost per find is the total search cost, not the lucky-run cost.
  • FFmpeg: several hundred runs at roughly $10,000 total found multiple vulnerabilities (H.264/H.265/AV1 codecs); three were fixed in FFmpeg 8.1.
  • Turning a known (N-day) issue into a working root exploit cost under $1,000 (half a day) in one case, under $2,000 (under a day) for a harder chained exploit.
  • Pricing: Mythos Preview is $25 / $125 per million input/output tokens — about 5–6× Claude Opus. Anthropic is subsidizing early testers via ~$100M in usage credits (so Palo Alto didn’t actually pay the $1M).
  • Cost routing: Palo Alto cut spend by having Mythos plan the attack and a cheaper model (e.g. Opus 4.7) execute it.
  • Mythos finds subtle, long-lived bugs — the oldest a 27-year-old OpenBSD flaw and a 16-year-old FFmpeg bug that scanners had tested millions of times. Anthropic reports “thousands” of high-severity zero-days across every major OS and browser; per Gartner, under 1% have been patched.

Sources: https://red.anthropic.com/2026/mythos-preview/ · https://anthropic.com/glasswing · https://www.informationweek.com/cybersecurity/anthropic-s-mythos-forces-a-rethink-of-vulnerability-management

short vs long horizon grades

explosion condition

Explosion condition: if each 1% increase in effective compute can find a >1% win in algorithmic efficiency.

If this holds then you can train a series of models, each one discovers new algorithmic efficiencies, which makes the successive model better, so that algorithmic efficiency grows at a super-exponential rate.

Consider a simplified case with fixed model training compute \(T\), and algorithmic efficiency \(A\). Suppose that, each generation, we run a model to optimize algorithms for the next generation’s model, \(A_{t+1}=f(A_tT)\), where \(f(.)\) . Then \(A\) will grow at a super-exponential rate if and only if \(f()\) has an elasticity greater than 1. We could use our optimization experiments to map out the function \(f(\cdot)\).

new graphs

Log axes

misc

Tom Davidson question

Q: How is this different from CES task model, where agents can only do some things?

A: Great question - you could consider each apple, or horizontal slice, a "task", & then do a task model, a few differences: (1) the full model interprets effort as cumulative (so closer to a Jones model) rather than delivering separablee period-by-period utility (like a task model); (2) in our model it's constant-returns across slices, diminishing-returns within-slices, while task models are usually the other way around; (3) I *think* if you assume isoelastic returns then you won't get rabbit-vs-hare effects, i.e. to match this you need decreasing elasticity, which is what our functional form has.
Start with the simplest model of AI progress:
  1. There is a fixed mapping from training expenditure to general-purpose model intelligence.
  2. Training expenditure has been increasing by ≈5X per year.
  3. Training efficiency has been increasing by ≈5X per year due to R&D.[^estimates]

We can draw this as follows:

Humans have been painstakingly pushing those blue curves to the left. We now are seeing clear signs of agents joining the effort and contributing to that leftward movement, and we want to know what to expect.

Observations.

The figure below shows the equilibrium, holding \(h_t\) fixed. The red curve shows total apples harvested as a function of agent height. The blue line shows the apples needed to sustain a given height (from inverting the self-improvement equation at steady state: \(a = \bar{a}+(\lambda-\alpha)/\beta\)). Where the curve is above the line, surplus apples drive \(\lambda\) upward; where below, the deficit pulls it back. The equilibrium is at the intersection.

–>

<!–

Conditions

Agents contributing to AI research.
With \(\lambda_n = 0\), \(a_n = h_n \to 1\). So the agent can ever turn on iff \(\bar{a} < 1\). If \(\bar{a} \ge 1\), \(\lambda_n \equiv 0\) forever. Activation-time approximation: \(h_n = 1 - p^n \ge \bar{a}\) \(\Leftrightarrow\) \(n \ge \ln(1-\bar{a})/\ln p\); \(p\) mainly shifts when activation happens.
Getting to superhuman AI research.

As \(n \to \infty\), \(h_n \to 1\). If \(\lambda_n < 1\), \(a_n \to 1\); if \(\lambda_n \ge 1\), \(a_n = \lambda_n\). So asymptotically \(\lambda_{n+1} \to f(1)\) with \(f(1) = 0\) if \(1 < \bar{a}\), and \(f(1) = \alpha + \beta(1-\bar{a})\) if \(1 \ge \bar{a}\). So takeoff past human level (eventually \(\lambda_n > 1\)) iff \[\boxed{\alpha + \beta(1-\bar{a}) > 1.}\]

Interpretation: “If the orchard were fully human-level (\(a=1\)), would the next agent be at least human-level?” If not, the system stays below 1. This condition is essentially independent of \(p\).

Self-sustaining AI research.

For \(\lambda_n \ge 1\), \(a_n = \lambda_n\) and \[\lambda_{n+1} = \alpha + \beta(\lambda_n - \bar{a}).\]

  • Runaway / hard takeoff iff \(\boxed{\beta > 1}\) (roughly geometric growth in \(\lambda_n\)).

  • Soft takeoff / saturation iff \(\boxed{\beta < 1}\): convergence to \[\lambda^* = \frac{\alpha - \beta\bar{a}}{1-\beta}\] (provided the system crosses 1 first).

  • Knife-edge \(\beta = 1\): linear growth.


Variation: Non-uniform apple density

Now let the tree have a general shape.

Suppose apples have density \(f(x)\) over heights \(x \ge 0\), with cumulative mass \[F(\lambda) \equiv \int_0^\lambda f(x)\,dx,\] normalized so that \(F(1)=1\). This keeps the interpretation that a human can reach one unit of cumulative progress, but allows the orchard to be bottom-heavy or top-heavy above that point.

The apples harvested by period \(n\) are now \[a_n = F(\lambda_n) + \big(1-F(\lambda_n)\big)_+\,h_n.\]

This is the same logic as before: the agent gets everything below its reach, while humans contribute a fraction \(h_n\) of the remaining human-reachable mass.

Implications.

The conditions for turning on and for crossing the human frontier are unchanged:

  1. The agent can activate iff \(\bar a < 1\).
  2. The agent eventually gets past human level iff \[\alpha + \beta(1-\bar a) > 1.\]

The reason is that below human level only the total mass up to height \(1\) matters, and we normalized that mass to 1.

Above human level, the shape of the tree matters.

Once \(\lambda_n \ge 1\), humans no longer contribute and the dynamics become \[\lambda_{n+1} = \alpha + \beta\big(F(\lambda_n)-\bar a\big).\]

The local gain from an extra bit of reach is therefore \[\frac{d\lambda_{n+1}}{d\lambda_n} = \beta f(\lambda_n).\]

So there is a simple density threshold: \[\boxed{\text{local RSI at height }\lambda \text{ iff } \beta f(\lambda) > 1 \iff f(\lambda) > 1/\beta.}\]

Interpretation: an extra unit of reach at height \(\lambda\) exposes about \(f(\lambda)\) extra apples, which in turn generate \(\beta f(\lambda)\) extra reach next period. If that amplification factor exceeds 1, improvements compound rather than dying out.

A stronger condition for sustained unbounded takeoff is that for sufficiently large \(\lambda\), \[\alpha + \beta\big(F(\lambda)-\bar a\big) > \lambda.\] But the density threshold \(f(\lambda)>1/\beta\) is the clean local criterion: it tells us where the orchard is dense enough for recursive improvement to reinforce itself.

This suggests a concrete empirical question: how dense are the “high apples” in AI R&D? If hard ideas remain dense enough above the current frontier, then crossing human level could quickly become self-reinforcing. If density falls off fast with height, we should expect saturation instead.