The combined criterion requires a shared-bias gain and limited regression on clean observations.
A PUBLIC EXPERIMENT IN TRUST
More agreement.
Or better evidence?
When observations share an error, repetition can make the wrong answer feel certain. Explore when a separate source is worth its cost.
Source-aware confirmation · Research experiment 004
Interactive evidence edition 005
Loading the frozen experiment…
01 / FINDINGS, INCLUDING THE FAILURE
A useful mechanism.
An unfinished result.
These are the registered results from the frozen experiment. Exploring the records below does not change them.
Learned selector minus independent-bias system.
Maximum allowed mean regression: 0.010.
Loading the learned-versus-control comparison…
Intervals are paired 95% bootstrap intervals over 96 episode clusters. The three learned seeds are averaged within each episode; the strongest control is reselected in each bootstrap replicate.
02 / FOLLOW THE EVIDENCE
Same world.
Different decisions.
This is a replay of recorded experiments. Choose any retained episode and inspect what each policy saw, bought and predicted. No new inference runs in your browser.
Choose an environment after the evidence loads.
The price of checking.
| Policy | Accuracy | Paid cost | Utility |
|---|
Accuracy is weighted by the episode’s target priorities. Utility is weighted accuracy minus paid observation cost, not a percentage. Learned results here use the selected checkpoint.
Fixed traces only: this changes the price used to score recorded actions. Policies do not adapt or choose new observations. It does not change the registered benchmark above.
Keep the wider view.
Utility at the selected post-hoc price. Learned bars average all three checkpoints across the same 96 worlds. These are descriptive comparisons, not new evaluation results.
See the trail behind the answer.
Recorded prediction endpoints
“Initial” means after the four inherited reports. “Final” means after all recorded actions. Intermediate group predictions were not retained and are not reconstructed.
Scoring truth comes from the synthetic world. The acquisition policy did not receive these hidden labels.
Record identity and source binding
03 / WHAT THIS SYSTEM ACTUALLY DOES
Small enough to inspect.
Precise enough to challenge.
A bounded research system for acquiring evidence. It does not establish AGI, open-domain factual reliability or a new foundation language model.
A supplied world model
Four target groups, sixteen labels per group, and an explicit shared source-bias variable. An exact Bayesian predictor represents the stated assumptions.
A learned decision to check
Three 26,792-parameter selectors learn acquisition utilities from raw observation history. The predictor and acquisition teacher remain programmed.
A real cost to confirmation
Every policy inherits four cheap reports, then receives four resource units. A cheap observation costs one unit; a reference costs two. Separate sources can still fail.
04 / STUDY, ADAPT, ATTRIBUTE
Ideas have a lineage.
These public sources inform our questions and design choices. Their methods and benchmark claims were not reproduced here. No upstream code or weights were imported.
What makes a verifier useful?
Process reward model lessons: separate final-answer selection from locating errors.
↗DEEPSEEK-AI · 2025Check the checker’s explanation.
DeepSeekMath-V2: a favorable score does not guarantee a faithful account of the fault.
↗UKRAINIAN CATHOLIC UNIVERSITY · 2026Locate the failure mechanism.
Layer-resolved diagnostics: retrieving content and using it correctly are separate problems.
↗WEST UKRAINIAN NATIONAL UNIVERSITY · 2024Keep a clean counterpart.
Entity embellishment research motivates explicit corruption provenance and bounded checks.
↗OXFORD OATML · 2023Meaning, beyond repeated words.
Semantic uncertainty motivates distinguishing agreement from independent evidence.
↗OXFORD STATISTICS · 2021Choose the next observation.
Deep Adaptive Design provides prior art for learned sequential experiment selection.
↗Design transfers above are our inferences. The release also documents public CIA, NSA, IARPA and DARPA material, with source-specific retrieval and reuse limits. This is a selective research map, not an institutional ranking or endorsement.
05 / NOTHING HIDDEN BEHIND THE CHART
Take the evidence with you.
Inspect the source, download the frozen research build, and reproduce it locally. The public explorer replays its retained records.
Release and evidence fingerprints
Hashes detect byte changes relative to the retained catalog. They are not signatures or proof that the underlying observations are true.