An independent check of whether TORAX, Google DeepMind's open-source tokamak transport simulator, reproduces its own committed reference outputs on hardware its CI does not run on. Each entry's expectations were written and hash-sealed before its graded runs, every comparison had negative controls, and the misses are reported alongside the passes.
| Entry | Question | Result | Pre-registered expectations |
|---|---|---|---|
| E001 | Does the flagship ITER-hybrid rampup case reproduce the maintainers' reference on independent hardware? | Reproduced All 201 output variables within relative 1e-9. Worst: 1.25e-11. Two identical runs are bit-for-bit equal. Upstream sim suite: 62/64 pass (see E002). |
6 of 7 held. Missed: embedded run config was predicted equal apart from the version stamp. v1.4.3 records 3 extra default-valued settings. |
| E002 | Why do 2 upstream restart tests fail here? | Explained · addressed on main A change of one or two ulps in stored thermal energy (about 2e-16 relative) in heating-off cases gives dW/dt = 2.5e-7 W against an exact-zero reference. The test's absolute tolerance, 1e-8, sits below that round-off. Current main uses 1e-6, and all 4 restart tests pass there. |
Reproduced on macOS and Linux arm64 with both release-era and current dependencies. It does not occur on x86 (E005). E004 found the same artifact in a reference written on the maintainers' platform, so this is platform round-off, not an arm64 defect. |
| E003 | Does the full output of every referenced test case reproduce? (v1.4.3, 56 cases) | Reproduced All 52 maintainer-tested cases: the 5 main profiles within 1e-9. Every other difference is round-off on near-zero values, a solver residual, or inside the maintainers' own per-case tolerance. 2 QuaLiKiz cases could not run. |
5 of 6 held. Missed: predicted at least 80% of cases within 1e-9 on every variable; observed 20/52. Every miss sits on near-zero or exact-zero values, a solver residual, or inside upstream tolerance. The classifier for this was written after seeing early results. |
| E004 | Does unreleased main reproduce its own references? (55 cases) | Reproduced on shared variables Upstream suite 63/63. 49/51 maintainer-tested cases within 1e-9 on the 5 main profiles. The other 2 (TGLF neural-network transport) are within 1.8e-6, inside the maintainers' 5e-6. Not validated: the new per-model transport outputs; 46/53 references predate them. |
4 of 6 held; the classifier was pre-registered this time. Missed for 3 cases: the two TGLF cases above, and one heating-off case where the reference holds the E002 artifact and arm64 gives exactly zero. |
| E005 | Do the x86_64 legs match? (GitHub runner, AMD EPYC 7763) | Matches Flagship within 1e-9 of the reference and of the Apple silicon output (worst cross-architecture difference 1.4e-11). Heating-off dW/dt is exactly zero. Upstream suite 64/64 on v1.4.3 with release-era dependencies. |
6 of 6 held. The first runner was terminated mid-run with no results seen; the repair was sealed before the rerun. |
| RAPTOR | The paper's benchmark against the RAPTOR code | Not evaluated No RAPTOR reference output was available locally, and none is in the TORAX repository. |
Not run. |
v1.4.3 and main at 17cc32fb, installed unmodified. For E002, dependencies were also pinned to the release date.On both the release and main, every simulation case the maintainers compare against a reference reproduces within their own tolerances. The one exception is the pair of v1.4.3 restart tests covered below. Where the all-variable check misses, the cause is round-off on values that are zero or nearly zero, a solver residual, or a difference inside upstream tolerance.
On v1.4.3, two restart tests fail on arm64 because their absolute tolerance (1e-8) is smaller than ulp-level noise in a quantity that should be exactly zero. On x86 the same tests pass and the quantity stays exactly zero. Main raised the tolerance to 1e-6 in 9274e5c2 (after an earlier change in PR 2351), and the tests pass there. No issue was filed because the fix already exists.
Main now writes per-model transport outputs, and 46 of 53 committed references do not contain them. The maintainers' tests check 5 profiles on those cases, so nothing fails. It does mean those new outputs currently have no reference to reproduce against.
The scripts, sealed pre-registrations, comparison reports and test logs are in the GitHub repository. Anyone can rerun E001 and check these hashes with shasum -a 256 -c. Local paths were stripped from logs for publication; the sealed files are unchanged.
| Item | SHA-256 | Sealed (UTC) |
|---|---|---|
| E001 pre-registration | ab9bfb5bdadf502c724bc80c0ffe0124ca22047f2c7848a668ad284b0d17c8f0 | 2026-09-25 06:28 |
| E002 pre-registration | e739c2481e6c25ae590fed5738ee4aea258f246b46cef434e4fc359c8f76a725 | 2026-09-25 06:35 |
| E002 addendum (x86 via Rosetta) | 6d8c1b012bb04b11c5c9a1de8f6bd9e2fb273637a0c5a7bc632ee00bad14e515 | before that run |
| E003 pre-registration | 5301c14e02914d81293ac46a6f373b751e4a1505f2e2624d8f8e04872cc76205 | 2026-09-25 11:29 |
| E004 pre-registration | e898217d9378d9312a8a8ca99eb30260b9818655d2e5d38826585da6b2da59b0 | 2026-09-25 11:40 |
| E005 pre-registration | 38c5aa327d545183ccf9f5cdae56483435c839c5bdb011e62aad1bd4e45fd509 | 2026-09-26 02:46 |
| E005 addendum (runner repair) | 1d5414506ada5e4c731f794eead4bd1a1d75428128b41743873d852f68daf6ed | 2026-09-26 03:10 |
| Comparator (compare.py) | 05eeb7961341429836d75269bbaa52e689a118aac57d0a99d6dccd7f9fa3b99e | unchanged E001–E005 |
| Runner (run_case.py) | 5ea12e947234e8088feb56b8dd95f30da3259a371bf2345d84c16978540b28c4 | unchanged E001–E005 |
| Repository manifest at commit 58d8973 (454 files) | e5875fb4962e085db8024ff68a92b49ed9ebb83e74ac5355f0766af59b01c12a | 2026-09-26 |