Reproduction ledger · Tokamak transport simulation

TORAX Reproduction Ledger

An independent check of whether TORAX, Google DeepMind's open-source tokamak transport simulator, reproduces its own committed reference outputs on hardware its CI does not run on. Each entry's expectations were written and hash-sealed before its graded runs, every comparison had negative controls, and the misses are reported alongside the passes.

Release tested
v1.4.3 · 4aea2377
Development tested
main · 17cc32fb (2026-09-24, unreleased)
Platforms
Apple M1 Max (arm64, macOS); GitHub runner, AMD EPYC 7763 (x86_64). JAX CPU, float64
Runs · published
2026-09-25 to 26 · 2026-09-26
Author
Adem Vessell
AI assistance
Runs executed by Claude Code, an AI coding agent, under the author's direction. Codex reviewed the ledger and the maintainer note before publication, without rerunning the experiments.
Results

Five entries, with the misses shown

EntryQuestionResultPre-registered expectations
E001 Does the flagship ITER-hybrid rampup case reproduce the maintainers' reference on independent hardware? Reproduced
All 201 output variables within relative 1e-9. Worst: 1.25e-11. Two identical runs are bit-for-bit equal. Upstream sim suite: 62/64 pass (see E002).
6 of 7 held. Missed: embedded run config was predicted equal apart from the version stamp. v1.4.3 records 3 extra default-valued settings.
E002 Why do 2 upstream restart tests fail here? Explained · addressed on main
A change of one or two ulps in stored thermal energy (about 2e-16 relative) in heating-off cases gives dW/dt = 2.5e-7 W against an exact-zero reference. The test's absolute tolerance, 1e-8, sits below that round-off. Current main uses 1e-6, and all 4 restart tests pass there.
Reproduced on macOS and Linux arm64 with both release-era and current dependencies. It does not occur on x86 (E005). E004 found the same artifact in a reference written on the maintainers' platform, so this is platform round-off, not an arm64 defect.
E003 Does the full output of every referenced test case reproduce? (v1.4.3, 56 cases) Reproduced
All 52 maintainer-tested cases: the 5 main profiles within 1e-9. Every other difference is round-off on near-zero values, a solver residual, or inside the maintainers' own per-case tolerance. 2 QuaLiKiz cases could not run.
5 of 6 held. Missed: predicted at least 80% of cases within 1e-9 on every variable; observed 20/52. Every miss sits on near-zero or exact-zero values, a solver residual, or inside upstream tolerance. The classifier for this was written after seeing early results.
E004 Does unreleased main reproduce its own references? (55 cases) Reproduced on shared variables
Upstream suite 63/63. 49/51 maintainer-tested cases within 1e-9 on the 5 main profiles. The other 2 (TGLF neural-network transport) are within 1.8e-6, inside the maintainers' 5e-6. Not validated: the new per-model transport outputs; 46/53 references predate them.
4 of 6 held; the classifier was pre-registered this time. Missed for 3 cases: the two TGLF cases above, and one heating-off case where the reference holds the E002 artifact and arm64 gives exactly zero.
E005 Do the x86_64 legs match? (GitHub runner, AMD EPYC 7763) Matches
Flagship within 1e-9 of the reference and of the Apple silicon output (worst cross-architecture difference 1.4e-11). Heating-off dW/dt is exactly zero. Upstream suite 64/64 on v1.4.3 with release-era dependencies.
6 of 6 held. The first runner was terminated mid-run with no results seen; the repair was sealed before the rerun.
RAPTOR The paper's benchmark against the RAPTOR code Not evaluated
No RAPTOR reference output was available locally, and none is in the TORAX repository.
Not run.
Scale

How big the differences are

Largest relative differences, log scale The E002 round-off in stored thermal energy is about 2.5e-16. The E001 worst variable is 1.25e-11. The maintainers' default tolerance is 1e-9. The E004 TGLF neural-network cases reach 1.8e-6 on the main profiles, inside the maintainers' 5e-6 tolerance for those cases. The plus one percent heating control changes electron temperature by up to 9.7e-3. 1e-16 1e-14 1e-12 1e-10 1e-8 1e-6 1e-4 1e-2 default test tolerance 1e-9 TGLF-NN test tolerance 5e-6 E002 round-off 2.5e-16 E001 worst variable 1.25e-11 TGLF-NN main profiles 1.8e-6 +1% heating control 9.7e-3
Largest relative difference from the committed reference, on a log scale. Each dot is a single worst value from the ledger. Dashed lines are the maintainers' own test tolerances. The dark dot is a deliberate +1% change in heating power, included so the scale shows what a real physics change looks like.
Method

How each entry was run

  1. Pin the code. Tagged release v1.4.3 and main at 17cc32fb, installed unmodified. For E002, dependencies were also pinned to the release date.
  2. Seal expectations first. Each entry's expected outcomes and claim limits were written and SHA-256 hashed before its graded runs (hashes below). Results are graded against those files, misses included.
  3. Compare the whole output. An independent comparator checks every output variable against the committed reference, not only the 5 profiles the upstream tests check. Each variable gets the strictest tier it meets: bit-for-bit, then relative 1e-9, 1e-6, 1e-3.
  4. Prove the comparator can fail. Controls: a repeat run must be bit-for-bit identical; a +1% heating change must be flagged; a mismatched reference must be rejected; an injected change of 1e-8 must be caught at the 1e-9 tier, and one of 1e-10 must not be overstated. All controls behaved as expected.
  5. Explain every miss against scale. Relative error is misleading near zero, so each miss is also measured against the variable's own peak value.
Findings

What came out of it

The release reproduces cleanly on Apple silicon and x86

On both the release and main, every simulation case the maintainers compare against a reference reproduces within their own tolerances. The one exception is the pair of v1.4.3 restart tests covered below. Where the all-variable check misses, the cause is round-off on values that are zero or nearly zero, a solver residual, or a difference inside upstream tolerance.

A restart-test tolerance was below round-off, and main already fixed it

On v1.4.3, two restart tests fail on arm64 because their absolute tolerance (1e-8) is smaller than ulp-level noise in a quantity that should be exactly zero. On x86 the same tests pass and the quantity stays exactly zero. Main raised the tolerance to 1e-6 in 9274e5c2 (after an earlier change in PR 2351), and the tests pass there. No issue was filed because the fix already exists.

Most references on main predate its new output layout

Main now writes per-model transport outputs, and 46 of 53 committed references do not contain them. The maintainers' tests check 5 profiles on those cases, so nothing fails. It does mean those new outputs currently have no reference to reproduce against.

Limits

What this does not show

Evidence

Sealed hashes

The scripts, sealed pre-registrations, comparison reports and test logs are in the GitHub repository. Anyone can rerun E001 and check these hashes with shasum -a 256 -c. Local paths were stripped from logs for publication; the sealed files are unchanged.

ItemSHA-256Sealed (UTC)
E001 pre-registrationab9bfb5bdadf502c724bc80c0ffe0124ca22047f2c7848a668ad284b0d17c8f02026-09-25 06:28
E002 pre-registratione739c2481e6c25ae590fed5738ee4aea258f246b46cef434e4fc359c8f76a7252026-09-25 06:35
E002 addendum (x86 via Rosetta)6d8c1b012bb04b11c5c9a1de8f6bd9e2fb273637a0c5a7bc632ee00bad14e515before that run
E003 pre-registration5301c14e02914d81293ac46a6f373b751e4a1505f2e2624d8f8e04872cc762052026-09-25 11:29
E004 pre-registratione898217d9378d9312a8a8ca99eb30260b9818655d2e5d38826585da6b2da59b02026-09-25 11:40
E005 pre-registration38c5aa327d545183ccf9f5cdae56483435c839c5bdb011e62aad1bd4e45fd5092026-09-26 02:46
E005 addendum (runner repair)1d5414506ada5e4c731f794eead4bd1a1d75428128b41743873d852f68daf6ed2026-09-26 03:10
Comparator (compare.py)05eeb7961341429836d75269bbaa52e689a118aac57d0a99d6dccd7f9fa3b99eunchanged E001–E005
Runner (run_case.py)5ea12e947234e8088feb56b8dd95f30da3259a371bf2345d84c16978540b28c4unchanged E001–E005
Repository manifest at commit 58d8973 (454 files)e5875fb4962e085db8024ff68a92b49ed9ebb83e74ac5355f0766af59b01c12a2026-09-26