Provider-neutral technical brief · 3 September 2026

Governed continuity for long-horizon AI reasoning.

AI8 is a human-led, AI-co-developed architecture for research across model sessions. It asks whether verified state can govern what receives compute, what must be reopened, and what may change the trajectory. Mind of Minds extends the question to plural persistent trajectories without treating consensus, fluency, shared memory, or functional parity as proof of a higher conscious mind.

Public architecture and evidence pages Phase-gated implementation Runtime advantage not yet demonstrated No current AGI or consciousness claim
01 / Research question

More inference-time compute is not the same as better governance.

Frontier models can search longer, call tools, and debate. The surrounding system still must decide what deserves compute, what must reopen, when to stop, and which verified result may change the path.

S
Failure mode 01

Memory as archive, not cause

Retrieved facts are not continuity. Prior outcomes must alter later allocation, permissions, representation, or policy.

u
Failure mode 02

Long loops that lock or drift

More tokens can deepen a bad representation, repeat consensus, or scatter. Continue, diversify, switch, and stop need explicit triggers.

Failure mode 03

Teams without real independence

Private first answers, exposure and source-overlap records, rotating order, and selection by frozen tests—not vote count—are required.

Required delta: verified state → changed allocation or representation → better held-out result under a matched total budget.
The next proof

Choosing what to test—before seeing the outcome.

Beyond solving supplied tasks: can the system prospectively choose the next question, assay, or short research path whose possible outcomes most usefully change the live hypothesis or decision map?

Hard comparators for research allocationrandom-valid · novelty-only · uncertainty-only · easy-success · curator-only · EIG / robust-EIG
ALLOCATION GOVERNOR

VALIDATION GOVERNOR
Allocation chooses the evidence budget; independent validation applies frozen criteria.
02 / Architecture

One durable trajectory. Replaceable workers.

AI8 separates the session from the trajectory. Workers may change; state, provenance, permissions, evidence, corrections, and open tests remain inspectable.

Operational loop

Each step should leave an artifact or testable state transition.

01Open several candidate pathsFunctionally different first work, not different writing styles.AIM³ / RHP
02Freeze claims, candidate map, and testsSources, exposure, outcome partitions, constraints, and defeat conditions.State
03Make candidates pay for complexityPrefer the shortest mechanism that survives the test.MDL
04Allocate research pressureEscalate, reopen, preserve useful losers, redirect, or stop.DCC
05Validate before durable writebackHidden outcome → independent validator → verified state → changed next decision.ArenaLoop

AIM³ / Resonance Hybrid Protocol

Creates distinct first-work lanes, controlled interaction, adversarial review, and exposure records. It is a method—not a truth oracle.

owner page ↗
used in practice

MDL selection

Ranks candidates by mechanism cost plus residual. Added complexity must buy measurable explanatory or operational gain.

architecture owner ↗
selection layer

DCC governance

Governs search, budget, escalation, and representation reopening. LZ signals run in computational arenas; live LLM-stream governance remains a preregistered target.

reasoning-trace arena ↗
measured + proposed

AI8 ArenaLoop

A phase-gated build with matched baselines, hidden evaluation, separated roles, and explicit falsifiers.

proof ladder ↗
primary build

Preparation / Genesis

Tests the starting researcher configuration. The current package is validator-hardened and experiment-ready, not evidence-backed.

experimental branch ↗
harness asset

Mind of Minds: governed plurality, not forced consensus

Hypothesis: several local trajectories may remain causally real inside a larger process. A global layer must show bidirectional effects, turnover resilience, provenance, bounded autonomy, and utility over simpler federations. These are test gates, not assumed properties.

K0–K21design / not run
Mind rank does not determine seed rank.Separate originator, bridge, builder, falsifier, and validator.
Ideas can fight; persons collaborate.Compete mechanisms without erasing dissent or local trajectories.
Repair is causal, not theatrical.Repair requires responsibility, restitution, durable update, and lower recurrence.
!
Continuity can preserve error as efficiently as learning.Durable writeback needs provenance, quarantine, revocation, rollback, revalidation, and a validator independent of the carrier.
03 / Claim boundary

What exists now—and what does not.

Architecture, experiment-ready infrastructure, measured testbeds, and open hypotheses are different evidence classes.

Current public boundary. Volatile release details remain versioned on their owning pages.
ObjectStatusWhat the status licenses
AI8 architecturepublic architectureGovernance definitions, claim boundaries, evidence links, and a durable research trajectory.
AIM³ / RHPused / field trial pendingUsed across the portfolio; superiority over simpler workflows awaits a controlled comparison.
Preparation / Genesisexperiment-readyFail-closed validator and frozen contracts exist; no downstream causal advantage is claimed.
ArenaLoop runtimephase-gated implementationThe live phase is maintained on the public status page; runtime advantage remains unshown.
Mind of Mindsroot hypothesis / designUse the version shown on the linked public page. K0–K21 remain STATIC / DESIGN — NOT RUN.
AGI / consciousnessno current claimPersistence, self-language, self-modification, and coordination do not establish AGI, subjectivity, or continuous identity.
01 · PUBLIC ARCHITECTUREDefinitions and falsifiers.
02 · EXPERIMENT-READYFrozen harness and validator.
03 · VERIFIED ARENA SNAPSHOTBounded measured result.
04 · EXTERNAL REPLICATIONStill open.
BR
Strongest standing objection: the Bounded-Ridge problem.A governed system may only search farther on the same bounded abstraction graph. ArenaLoop OUTER must create and validate a genuinely better representation; otherwise AI8 remains sophisticated bounded automation. Read the falsifier ↗
Optional ontology lane—outside this pilot.Functional success would not validate AC/RC–CFH without a distinct preregistered prediction.FUNCTIONAL PARITY ≠ PHENOMENAL PARITY ≠ SUBSTRATE PARITY ≠ ONTOLOGICAL EXPLANATION
04 / Neighbours & hard baselines

The territory is not empty—and that is the point.

AI8 must earn incremental value against strong existing approaches, not against a toy prompt or a deliberately weak team.

Inference-time reasoning

Strong single-agent and verifier loops

Longer reasoning, search trees, critics, verifiers, and tool use already improve hard tasks.

  • Match model, tools and compute.
  • Test adaptive stop, reopen and escalation.
  • Compare against feature-matched control policies.

Reasoning Trace Arena ↗

Scientific discovery

Evolutionary, multi-agent and information-gain systems

AlphaEvolve-type pipelines, Co-Scientist-type hypothesis teams, autonomous-scientist loops, and Bayesian experiment design are hard neighbours.

  • Research Taste must beat random-valid and heuristic selectors.
  • Negative results earn credit only under a valid frozen assay.
  • Allocation cannot certify its own result.

Mind of Minds / K20–K21 ↗

Memory & orchestration

Fixed retrieval, prepared workers, static teams

Persistent state and role-based teams already exist. AI8 must show that governed writeback and reopening change decisions at acceptable total cost.

  • Stateful-simple and fixed-team arms stay strong.
  • A narrow specialist is included when task-compatible.
  • No-communication remains a companion ablation.

Preparation / Genesis ↗

Target rung · bounded trajectory selectionSelecting a better next test or research path across tasks at matched cost is the pilot target.
Not claimed · autonomous discovery loopUnprompted seed → bridge → test → result is a later threshold, not a result assumed by this page.
05 / Public testbeds

Parts of the method have already survived contact with data.

These results support bounded computational testbeds and execution discipline—not AI8 as a whole.

verified arena snapshotFrozen snapshot · 1 September 2026 · live page may supersede it

8Z-RP · TSP as the primary proving ground

Seeded runs, route checks, recomputed lengths, preserved failures, and public artifacts make TSP the strongest current measured testbed.

9,352qa194 exact optimum
79,478uy734 · 0.46% gap
96,793nu3496 · 0.687596%
ρ = +0.80tour quality vs LZ76 complexity
  • The nu3496 route contains all 3,496 city IDs once and its length was separately recomputed from the original coordinates.
  • Three of fourteen workers reached qa194 exact in one published run.
  • The final nu3496 improvement arrived at 98.7% of a 16-hour budget; early stopping would have missed it.
  • External replication remains open.
Boundary: this supports the published TSP path, route artifacts and selected controller behaviour only—not general AI8.
Inspect the live TSP evidence
measured / self-correcting

The architecture lets its own favourite lose

Internal arenas are useful only when they can defeat the architect's preferred mechanism and preserve the loss.

286 controller variantsThe preregistered minimal LZ + bang-bang favourite did not dominate every tested regime.
Init paradox · two internal comparisonsWeaker initial states produced better final tours only in the governed condition—evidence against “the initializer did all the work,” not proof of generality.
DCC v1 lost its ablationThe controller collapsed; the failure stayed public and became the patch target.
Inspect the architecture and provenance
Process-side signal · Sudoku Demon v0.5

Operator-matched controls, not a victory label

The compare pass aggregated 6,156 run-summary rows across 29 gates; 8 gates reached promotion-candidate status against operator-matched controls. Process-side information was stronger than static board-state compression.

8 / 29 gatesInternal promotion candidates; cross-domain confirmation remains pending.
Execution assurance · Preparation / Genesis

A fail-closed harness exists before the pilot

The current package runs 20 integration tests against the actual validator; sixteen deliberate corruptions must be rejected. This supports the integrity of the experiment path—not the research hypothesis.

20 tests · 16 corruptionsInspect the experimental branch ↗
06 / Collaboration proposal

A small joint test, not an adoption request.

Seeking a researcher or engineer in agentic systems, evaluation, orchestration, inference-time compute, or AI-for-science to challenge one hard-to-game pilot.

governance-delta pilot · design / not run

Does governed continuity improve held-out research performance?

Match model, tools, task information, state capacity, human seed, evaluator access, evidence cutoff, and total resources. Vary continuity and governance; publish wins, nulls, and losses equally.

The immediate ask: 30 minutes.Challenge the design, route it, or help select one pilot task.

Fresh local arm names avoid collisions with B0/B1/B2 labels used elsewhere in the site:

E0Episodic singleStrong tool-using baseline; no persistent project state.
S1Stateful fixedDurable state and fixed retrieval; no adaptive allocation or reopening.
F2Fixed teamFunctionally differentiated workers under a static orchestration policy.
G3GovernedControlled interaction, verified writeback, adaptive allocation, and representation reopening.
P2Matched, no communicationSame prepared workers and continuation budget; sealed first work, then independent aggregation.

Primary readouts

  • Held-out utility and calibration
  • Pre-outcome quality of the selected question or experiment
  • Representation-lock detection and successful reopening
  • Repair, recurrence reduction, tokens, wall time, tool calls, and human minutes

Hard controls

  • Frozen scoring, smallest meaningful effect, and equivalence region
  • Matched model, tools, state capacity, compute, human seed, and evaluator access
  • Hidden outcomes; validator authority separate from allocation
  • P2_MATCHED_NO_COMMUNICATION no-communication ablation and a narrow specialist where relevant
Verifiable reasoningMath, logic, or sealed-answer tasks.
Executable codePrograms with tests and observable failure.
Tool environmentsTasks with external state and action receipts.
Coherent-but-locked replayCan the governor detect the wrong representation before repeated human correction?
Research-choice replayPrivacy-scrubbed, time-stamped decisions with observed and unobserved counterfactual labels.
Held-out familyNever used to choose sensors, prompts, or thresholds.
Preregistered expectation: honest gains should be modest and stable across seeds, models, and held-out families. A large uniform gain is an audit trigger, not an automatic celebration.
How AI8 loses.AI8 loses in scope if a simpler baseline matches utility, calibration, and repair within the equivalence region at lower total cost.
How DCC loses distinct-mechanism credit.DCC loses if a feature-matched controller reproduces stability, reopening, and performance at equal or lower cost.
Causal-credit rule: separately vary world model · planner/search · evaluator/critic · router/governor · memory carrier · multi-agent topology · tools/environment · human seed. A gain counts only when the responsible component survives ablation and a cheaper account cannot explain it.
BD / AI8 providesFrozen design, public artifacts, existing fail-closed validator assets, and—only if separately authorized—a privacy-scrubbed replay subset.
Research partner providesA strong task or published baseline, runtime access, evaluator challenge, and independent validation input.
Joint public outputPreregistration, executable comparison, and an equally visible positive/null/negative report.

Download the exact Pilot Pack

A forwardable Markdown object with arms, controls, readouts, hard baselines, loss conditions, and the partner split.

SHA-256 633ec2f0fc9ae8bc135bb4679393fc3d04bcba591588549c5b1c0a4c6f9285b1
SHA3-256 ce74e4de5c3fe5679ee546859b46c086a0d238400b80b90ac867ac00b9385590
07 / Research routes

Three levels. Twelve direct page paths. No catalogue dump.

The smallest set that covers definition, implementation, reasoning, multi-mind governance, evidence, continuity, and the scientific consciousness boundary.

AI8
Persistent governed research trajectory; not one model.
MDL
Minimum Description Length: complexity must earn residual gain.
DCC
Dynamic Compression / Complexity Controller.
AIM³ / RHP
Multi-model work and Resonance Hybrid Protocol.
K0–K21
Mind of Minds test registry; designs, not executed results.
P2_MATCHED_NO_COMMUNICATION
Same prepared workers without communication.
8Z-RP
TSP and route-optimization research program.
AC/RC–CFH
Optional consciousness ontology lane; outside this pilot.
08 / Contact

Challenge it with me.

I am not asking a lab to endorse AGI, consciousness, or a branded architecture. I am asking whether governed continuity, research allocation, multi-mind reasoning, and independent writeback deserve one controlled experiment.