SZLFOUNDATIONA11oy ↗

A PUBLIC EXPERIMENT IN TRUST

More agreement.
Or better evidence?

When observations share an error, repetition can make the wrong answer feel certain. Explore when a separate source is worth its cost.

Source-aware confirmation · Research experiment 004
Interactive evidence edition 005

5,184recorded policy rollouts
576distinct synthetic worlds
7policy types compared
3trained acquisition models

Loading the frozen experiment…

01 / FINDINGS, INCLUDING THE FAILURE

A useful mechanism.
An unfinished result.

These are the registered results from the frozen experiment. Exploring the records below does not change them.

OVERALL REGISTERED GATELoading

The combined criterion requires a shared-bias gain and limited regression on clean observations.

SHARED-BIAS UTILITY GAIN—Awaiting evidence

Learned selector minus independent-bias system.

CLEAN-SENSOR UTILITY CHANGE—Awaiting evidence

Maximum allowed mean regression: 0.010.

THE STRONGEST CONTROL MATTERS

Loading the learned-versus-control comparison…

Intervals are paired 95% bootstrap intervals over 96 episode clusters. The three learned seeds are averaged within each episode; the strongest control is reselected in each bootstrap replicate.

02 / FOLLOW THE EVIDENCE

Same world.
Different decisions.

This is a replay of recorded experiments. Choose any retained episode and inspect what each policy saw, bought and predicted. No new inference runs in your browser.

Choose an environment after the evidence loads.

ALL POLICIES / ONE EPISODE

The price of checking.

Recorded outcomes
All seven policies on the selected episode
PolicyAccuracyPaid costUtility

Accuracy is weighted by the episode’s target priorities. Utility is weighted accuracy minus paid observation cost, not a percentage. Learned results here use the selected checkpoint.

1.00×
Free evidence3× price

Fixed traces only: this changes the price used to score recorded actions. Policies do not adapt or choose new observations. It does not change the registered benchmark above.

ALL 96 EPISODES / THIS ENVIRONMENT

Keep the wider view.

Utility at the selected post-hoc price. Learned bars average all three checkpoints across the same 96 worlds. These are descriptive comparisons, not new evaluation results.

OBSERVATION REPLAY

See the trail behind the answer.

0 / 0

REPLAYED OBSERVATION—
INCREMENTAL PAID COST—
RESOURCE UNITS LEFT—
RECORDED P(SOURCE BIAS = 0)—

    Recorded prediction endpoints

    “Initial” means after the four inherited reports. “Final” means after all recorded actions. Intermediate group predictions were not retained and are not reconstructed.

    Scoring truth comes from the synthetic world. The acquisition policy did not receive these hidden labels.

    Record identity and source binding

    03 / WHAT THIS SYSTEM ACTUALLY DOES

    Small enough to inspect.
    Precise enough to challenge.

    A bounded research system for acquiring evidence. It does not establish AGI, open-domain factual reliability or a new foundation language model.

    01

    A supplied world model

    Four target groups, sixteen labels per group, and an explicit shared source-bias variable. An exact Bayesian predictor represents the stated assumptions.

    02

    A learned decision to check

    Three 26,792-parameter selectors learn acquisition utilities from raw observation history. The predictor and acquisition teacher remain programmed.

    03

    A real cost to confirmation

    Every policy inherits four cheap reports, then receives four resource units. A cheap observation costs one unit; a reference costs two. Separate sources can still fail.

    THE NEXT TESTABLE QUESTION

    Can a selector avoid unnecessary confirmation while staying robust when observations are independently corrupted?

    That requires a new registered protocol and fresh evaluation worlds. The exposed v0.4 challenge set cannot become a new held-out test.

    05 / NOTHING HIDDEN BEHIND THE CHART

    Take the evidence with you.

    Inspect the source, download the frozen research build, and reproduce it locally. The public explorer replays its retained records.

    Release and evidence fingerprints

    Hashes detect byte changes relative to the retained catalog. They are not signatures or proof that the underlying observations are true.