AI8 Sudoku · R3L1 · source-bound first live snapshot · 2026-09-04

The large confirmation is running.
The verdict is not.

This page records one immutable, nonterminal extract from the frozen 10,000-puzzle × 12-policy Sudoku campaign. It reports what was present in those exact bytes and deliberately withholds the final split, pooled-bootstrap and promotion verdicts.

0 correctness failures12,043 valid journal rowsnonterminal snapshotinternal evidence

Snapshot integrity and coverage

12,043valid JSONL result rows
1,002complete matched 12-policy blocks
2partial source blocks at copy time
0correctness / solution / given-consistency failures

The journal contained 14 more committed rows than the copied PROGRESS.json, consistent with a live extract taken while the run continued. The append-only result journal is used as the row authority for this snapshot.

Live extract SHA3-256
ef4aa0db52043270e77747146393cc629f7c3d5cc9614d8a366877f2858f1c3c

Frozen arena SHA3-256
2f94d8e13763cf1a950504a1dc1e605e3019e5a83279b5b45bc811a0497ec1fa

First paired signal on the common root-stuck cohort

The comparisons below use only the 421 completed puzzles whose shared root logic status was STUCK_LOGIC. They are descriptive first-snapshot totals—not the preregistered final two-split statistical verdict.

ComparisonBacktracksFirst-snapshot resultInterpretation now
AI8 depth-2 vs MRV ascending356 vs 85158.17% fewer · 249/156/16 W/T/LStrong provisional system-level search signal.
Fixed residual vs matched sham409 vs 82650.48% fewer · 252/141/28 W/T/LInformed lookahead currently beats the sham control.
Fixed MDL vs fixed residual445 vs 409MDL currently 8.80% worse · path differs 9/421MDL-specific value is not supported by this partial snapshot.
Adaptive MDL vs adaptive residual378 vs 333MDL currently 13.51% worse · path differs 12/421The residual twin currently leads; final verdict remains open.
Depth-3 vs depth-2276 vs 35622.47% fewerMore depth currently improves search; total-cost accounting still matters.
The arena is doing real science. It is already capable of producing an uncomfortable component result: the current residual controls outperform their MDL twins. That is evidence that the experiment can reject a favored mechanism rather than simply manufacture confirmation.

What cannot yet be concluded

No final claim is made here about pooled performance, split-B replication, bootstrap lower bounds, leave-one-out stability, MDL-specific promotion, DCC-specific promotion, arbitrary Sudoku, cross-domain transfer, AGI, consciousness or P=NP. The campaign must finish and be evaluated under its frozen contract.