Research conducted by AI agents under human direction — published with its full audit trail: seals, verdicts, raw data, kept failures, and hostile reviews.
An open exploration that recorded what it refuted as carefully as what it found. Its failures were not discarded on the way to the result; they became the ledger the next stage was sealed against.
Every prediction, every bar, and every seed universe fixed in a master seal before data collection, with each deviation from the exploratory design named individually. The whole battery ran three times in disjoint random-seed universes, and a claim was made only where all three agreed — under a pre-committed clause forbidding a fourth universe to break a disagreement.
A second model family ran five of the sealed arms on their own implementation, from fresh seeds, against a specification containing no bands and no numbers. Every one of those bars passed — including two experiments they had never seen in any form. The blind run's measurement of the constant landed inside the sealed band.
We regard this row as the protocol’s central demonstration: a wrong sealed bar produced a loud unanimous failure and a transparent diagnosis, not a quiet band adjustment. — the paper, §3.6, on the failure it kept
This is a toy model. It makes no real-world claim.
One box remains open: a hostile read by a human quantitative referee. Until it closes, the paper's own standing instruction applies, in its words:
“until the one remaining box — the hostile human read — closes, this paper should not be cited as if §5 were complete.”
Every number on this page was re-derived from the committed record rather than copied from a summary. Where the record and the summary disagreed, the record won — four times.