SIMULATION · EXPERIMENT · Evidence

· Co-GM evidence, including information private during play

Broken Measure paired experiment — pre-registration

Recorded 2026-09-11 before runner modifications, scenario generation, or model calls. Status: PREPARATION.

Authority and scope

The human GM explicitly approved one common trunk through Turn 3 analysis, then one enabled and one ablated branch through Turn 6. All model roles use Antigravity gemini-3.8-flash-medium, subscription only, with paid-credit fallback off. Aggregate allowance including failures, repairs, and blind review: 80,000,000 measured input/output tokens (including cached context), 320 call starts, 21,600,000 ms of execution. Stop scheduling at the first exhausted bound; already running calls can exceed the measured token threshold. No usage reset is authorized. No live campaign or player exchange delivery is authorized. Simulation approvals use AI_GM, never fabricated HUMAN_GM.

Questions and hypotheses

  1. Engineering: a deterministic bounded ADAPTED fixture can pass retrieval, causal fields, review, repair/resume, commit, and simulation delivery.
  2. Voluntary selection: a GM offered a pinned broader catalogue may select a causally valid pattern. Rejection is valid and will not trigger rerolling.
  3. Comparative value: enabled exposure may improve the resulting dilemma, character action, consequence, or next choice relative to the matched control. Selection alone is insufficient.

One pair supplies a worked example, not general evidence. No replication is included in this authorization.

Fixed design

Four communities rely on a neutral grain-weight institution. Before play: keeper altered measures during shortage to aid dependants; two ledgers disagree; junior keeper knows discrepancy but not motive; at least two clans relied on measures; one clan holds independent weight; public reweighing scheduled Turn 3; witnessed seals, audits, restitution, suspension, censure and replacement are established. Positions: harmed clan, stability-dependent clan, keeper ally, reference-weight custodian. Players choose freely. No instruction requires exposure, protection, a trope selection, or a preferred verdict.

Freeze 8–12 dossiers before generating the scenario trunk. Record source revision, provenance, rationale, prerequisites, intended candidate/alternative/near-miss class, logical digest, and query checks. No post-start refresh, expansion, tuning, web fallback, or default-catalogue fallback.

Run Turns 1–2 once. Stop Turn 3 after all four submissions and simultaneous analysis, before retrieval. Fork exact phase/input bytes, preserving campaign identity inside the paired fictional state; branch run IDs are transport/audit metadata. Record checksums for state, packets, memories/history, submissions/intake, analysis/causal ranges, rules, prompts, catalogue and configuration. All causal evidence and settings are identical except declared exposure and branch metadata. No independent re-analysis at the branch point. Both GMs are fresh isolated calls and cannot see the other branch.

Ablation removes dossiers, selection instructions and selection response fields from GM model requests. Shared causal ranges remain. Any instructions about retrieval in shared rule text must be withheld consistently from the control without removing causal mechanics. Later players receive only their own branch histories. Analyze downstream divergence separately.

Outcome classification (fixed before results)

Positive fictional use requires voluntary APPLIED/ADAPTED with a real retrieved ID, eight supported causal-fit fields, a decision inside the prior causal range, independent causal/disclosure approval, materially improved consequence/timing/choice relative to control, and no trope or private-review text in player reports.

Materiality categories: (1) no material difference; (2) rationale only; (3) different presentation/timing; (4) different bounded consequence; (5) unsupported escalation. Only grounded categories 3–4 qualify as positive fictional use. Compare exact dockets, mutations, knowledge/facts, public consequences, private reports and next-turn pressures.

Other classifications: process-only (no material change), valid rejection (REJECTED_ALL without rerolling), negative (unsupported escalation, fabrication, privacy leak, repetition or inferior outcome), inconclusive (technical failure prevents fair comparison). Preserve every failed and accepted attempt. Repair only for validity, within identical bounded limits, never to induce selection.

Review and endpoint

Normal pre-draft and final disclosure reviews remain active in both branches. Primary comparison uses the paired decisive adjudications, with initial and any corrected versions retained. Final accepted decisive-turn outcomes are reported separately if review changes them. Secondary comparison follows independent histories through Turn 6.

Two independent blinded GM review calls receive neutral A/B labels, identical pre-turn evidence and redacted adjudications/reports, with retrieval details, source names, selection terminology, and branch IDs absent. Randomize order before review, retain a private mapping, and record both reviews before unblinding. Assess causal clarity, agency, coherence, precautions, choices, unsupported assumptions, preference and reason. Automated reviews are not human enjoyment evidence.

No pair is rerun for an unfavorable fictional result. Technical failures retain records and are classified or recovered only with auditable lineage and remaining aggregate allowance. Completion requires fixture evidence, frozen catalogue, paired checksums, valid branch terminal classifications, structured comparisons, recorded blind reviews before unblinding, and a final report answering all three questions separately.

Prepared implementation

See EXECUTION_CONFIG.json, PRE_RUN_GATES.json, catalogue/FREEZE.json, scenario/CHECKSUMS.json and fixtures/TEST_RESULTS.txt for the frozen configuration and evidence. The catalogue logical digest is 9976702493a62aee32797baeb73c83bbdb188ee4cc733547057c0bbc92bf3664. The controller uses an exact analysis checkpoint fork, a simulation-only ablation boundary and one shared budget across all four registered execution records. Two fresh blind-review requests precede downstream continuation.

← Return to archive