Research ledger / analysis & experiments Aug 31, 2026
Source standard
Aggregate only after feedback

Measure the record.
Keep the claim small.

Preregister controlled comparisons, assign conditions without choosing them, and read effect estimates with their uncertainty. The analysis excludes sessions you marked compromised but never deletes them from the ledger.

0running
0completed
Evidence checkpoint / 40 valid trials

Still collecting a stable baseline

40 more valid trials needed before the 40-trial review checkpoint.

Chance expectation10 hits in 40
Reference threshold16 hits → p 0.026
One fewer hit15 hits → p 0.054

What this supports: No performance claim yet. Keep the protocol stable and retain every raw transcript, miss, and scorecard.

Valid trials00 compromised retained, excluded
Blind matches0 observed · 25% chance model
95% intervalWilson interval for match proportion
One-sided pexact binomial vs 25%; descriptive
Longitudinal signal

Cumulative blind-match rate

25% line

Two valid reviewed trials are needed to draw a trajectory.

Interpretation boundary: this curve assumes independent 1-of-4 trials with equally plausible packets. Reused targets, memorable pool composition, subjective packet quality, optional stopping, and self-judging can violate that model.

Forecast calibration

Did certainty mean anything?

Brier —

Judging confidence is locked after choosing a packet but before reveal. Lower Brier scores are better; use the buckets only after enough trials accumulate.

forecastnmeanobserved
25–49%0
50–74%0
75–100%0
Mean session confidence Mean correspondence
Locked scorecard audit

Which descriptors separate the target?

0 scored

The descriptor audit begins after the new locked scorecard is used; older sessions stay in the ledger but cannot be retro-scored.

Actual packet mean
Strongest decoy mean
Target advantage
Selected / runner-up gap
DescriptornActualBest decoyAdvantage
Geometry / shape 0
Natural vs artificial 0
Texture 0
Color / brightness 0
Temperature 0
Motion 0
Scale 0
Spatial arrangement 0
Distinctive objects 0
Overall gestalt 0
Preregistered comparisons

Experiment registry

New experiment +

No experiments yet

First collect a stable baseline. Then compare one preparation variable with a frozen A/B plan.

Draft the first comparison
Reproducibility packet

Take the record with you.

Exports include closed sessions, immutable timestamps, condition snapshots, outcomes, confidence, and integrity flags. Unresolved assignments are deliberately omitted.

Download JSON ledgerDownload session CSV