Measure the record.
Keep the claim small.
Preregister controlled comparisons, assign conditions without choosing them, and read effect estimates with their uncertainty. The analysis excludes sessions you marked compromised but never deletes them from the ledger.
Still collecting a stable baseline
40 more valid trials needed before the 40-trial review checkpoint.
What this supports: No performance claim yet. Keep the protocol stable and retain every raw transcript, miss, and scorecard.
Cumulative blind-match rate
Two valid reviewed trials are needed to draw a trajectory.
Interpretation boundary: this curve assumes independent 1-of-4 trials with equally plausible packets. Reused targets, memorable pool composition, subjective packet quality, optional stopping, and self-judging can violate that model.
Did certainty mean anything?
Judging confidence is locked after choosing a packet but before reveal. Lower Brier scores are better; use the buckets only after enough trials accumulate.
Which descriptors separate the target?
The descriptor audit begins after the new locked scorecard is used; older sessions stay in the ledger but cannot be retro-scored.
Experiment registry
No experiments yet
First collect a stable baseline. Then compare one preparation variable with a frozen A/B plan.
Draft the first comparison ↗Take the record with you.
Exports include closed sessions, immutable timestamps, condition snapshots, outcomes, confidence, and integrity flags. Unresolved assignments are deliberately omitted.