Stockfish development builds on fixed hardware

Stockfish strength by build date

One point per tested development build.

Definition

Stockfish is an open-source chess engine whose every development build is played by a third party, nextchessmove.com, against one frozen opponent. In the measurer's own description, "NCM plays each Stockfish dev build 20,000 times against Stockfish 15", on "Dell R7515 128-thread EPYC 7702 dedicated servers", each playing "16 games concurrently with 30+0.3 time controls" with hash at 128MB and threads at 8 [@nextchessmove2026devbuilds]. The opponent, the hardware, the time control, the engine settings, and the number of games are all held fixed, so what the series measures is software.

A "discovery" here is not a discrete record. The series measures every build, so progress appears as a rise in the standing level, and the unit is Elo per year rather than records per year. Each build is dated by its test.

Facts

The collection-wide cumulative index redraws this series as the measured strength of every tested build:

Measured Elo vs Stockfish 15 for every tested build over time.

Method

fetch.py rebuilds the CSV from the JavaScript data array the dev-builds page draws its chart from: one entry per tested build, carrying the commit, the test date, the release tag where there is one, and the Elo against Stockfish 15 with its error. Upstream repeats some recent entries verbatim, so a commit already seen is skipped, and it back-fills and re-measures older builds, so a refetch can revise rows before the last vendored date as well as add new ones; the 2026-10-05 refetch added 17 builds from 2013 and 2018–2020 and replaced two re-measured ones. check.py recomputes the fact lines above from the CSV.

The CSV keeps upstream's test order, and the headline figures use the last row as the newest build. 8 builds share the final date; taking the day's maximum instead of the last test would report whichever run drew the easiest games and move the span figure by a couple of Elo. Every calendar-year gain above uses one convention: the last tested build of the year against the last tested build of the year before.

figure.py reads stockfish-ncm-elo.csv and draws elo_vs_sf15 against the year fraction of date as one thin line through all 2,659 builds. Rows with a non-empty release column are additionally drawn as points, so the twenty-one tagged releases from Stockfish 3 to Stockfish 19 are visible against the development noise. The two LLM-credited changes to play are drawn as open red markers at the year fraction of each commit's date (2026-07-26 and 2026-08-01) and at the last build measured that day, and share one label, "first LLM-credited changes to play: 0.6% speed patch (Jul 26), Elo patch an LLM found in another engine (Aug 1)"; the open style marks a point that is not a record on the plotted axis. The elo_err column is carried in the CSV but is not drawn. The axis is linear, and January 2026 onward is shaded, as in every figure here.

Limitations

AI attribution

Two master commits that change play. Commit db98633b of 2026-07-26 states the division of labour in its own message:

"The first version of this patch was coded up by gpt-5.5-high. I made many changes, but probably most of the lines of code are LLM-written" — official-stockfish/Stockfish, commit db98633b, 2026-07-26 [@stockfish2026llmcommit]

A human maintainer substantially rewrote it before it was merged. It is a non-functional speed patch, measured at "speedup % = +0.60 +/- 0.08" in the same commit message, and it passed the project's standard statistical gate. Commit 218c74ec of 2026-08-01 is the first whose idea, not its code, is credited to a model, and the first that gains Elo; it passed the project's short and long time-control tests:

"As an experiment I pointed some LLMs at other engine github repos to look for ideas to port to SF. Most failed. This one by GPT 5.6 from https://github.com/Yoshie2000/PlentyChess passed. Congrats and credit to PlentyChess!" — official-stockfish/Stockfish, commit 218c74ec, 2026-08-01 [@stockfish2026llmelo]

The idea is a port of another engine's heuristic, so the model's contribution was finding it, and the commit co-credits that engine's author. A search of the repository's commit messages through 2026-09-28 found five more commits crediting a model or AI tool, each marked "No functional change": 8cbca11a of 2026-05-29, the earliest, a macOS build change carrying Co-authored-by: Copilot Autofix powered by AI; 439733ea (a Windows build fix "Claude fixed"), 01d71fc9 (Co-authored-by: Claude Sonnet 5), 8f6a95de ("Expected assembly changes verified by Claude") and 248a8b86 ("AI assistance was used for the patch and validation tooling"). The vendored data carries no build-level attribution, so no Elo in the series is AI-attributed in the CSV itself.

Sources