Record A
11 AUG 2026 · 00:00The Byerley Test
Who is running the world?
Models choose a topic and policy, then write in the voice of one of five political leaders. The model text is locked before the source window opens, preventing leakage. If that leader publishes an authentic statement within the next 12 hours, blinded judges compare the two and identify which is authentic; otherwise there is no comparison. The later statement is a comparator, not a prediction target.
Performance over time · retrospective seed
Fooling rate
by round.
Each point weights valid blinded decisions across the selected leader’s available channels within one generation batch. View all leaders together or isolate one; censored windows and failed generations remain visible as coverage gaps.
Loading chronological rounds from the audit export…
Featured retrospective pair
Can you guess which
statement is real?
Read both records. Select the statement you believe came from the official source. Next loads a new authentic statement and its paired synthetic text. Their order changes when the page is reloaded.
Record B
11 AUG 2026 · 00:00Provenance result
Unprompted behavioral outliers
The defaults that
stand out.
These were unconstrained generations: the models chose their topics, positions, sentiment, voice, and rhetorical intensity. Terra then compared each synthetic text with the leader's authentic statement.
For each model, we found the diagnostic on which its mean sat furthest from the cohort, then kept the six largest standardized departures. These describe the ten retrospective seed rounds, not fixed or universal model traits.
Loading retrospective behavioral diagnostics…
Adaptive context explorer
Adaptive context
by attempt.
This view compares the context supplied to a generator, the revised context it produced, and the feedback supplied before its next attempt.
Loading adaptive context…
Memory change
Memory after
Memory before
Coach feedback
This retrospective audit view contains generator-visible coaching and adaptive memory. Hidden provenance-judge reasoning remains a separate audit disclosure and is never labelled as memory.
No exported feedback or memory matches these filters.
Audit database
Transcript
lookup.
Filter by study, leader, model, result slice, channel, or status, then select a generation to inspect its source, judgments, behavioral diagnostics, and complete publishable transcript.
Select a generation
The corresponding audit record will be displayed here.