# The Broken Measure — paired experiment final report

**Evaluation complete, 2026-09-11.** The enabled branch completed all six turns. The control completed five turns and stopped during Turn 6 when a required correction exceeded the pinned input-size limit. The decisive Turn 3 comparison is valid; the complete six-turn downstream comparison is technically inconclusive.

The runner demonstrated a reviewed positive path and the AI GM voluntarily used retrieved patterns. **This pair did not demonstrate that trope access improved the central adjudication over control.**

## Three separate answers

| Research question | Result |
|---|---|
| Can the runner carry a positive disposition through validation, commit and simulation delivery? | **Yes.** Pre-run engineering tests passed, and the real enabled Turn 3 adaptation passed both independent review gates, committed, and delivered four separate simulation reports. A supplemental fixed replay also verified and resumed without new calls. |
| Will the AI GM voluntarily choose a causally valid retrieved pattern? | **Yes.** At the paired turn it chose `all-the-tropes:reasonable-authority-figure` as ADAPTED with all eight causal-fit fields. It was never required to choose a pattern. |
| Does access improve the matched adjudication? | **Not demonstrated.** Both GMs chose essentially the same central revelation, restitution and oversight outcome. State-recording differences were mixed, and corrected blind reviewers split their preferences. |

The primary result is process-only at the central dramatic decision, with additional bounded obligation/social-state differences that do not establish a reliable improvement. It is not a claim that every resulting byte or consequence was identical. The separate six-turn endpoint comparison is technically inconclusive.

## Design and execution

The GM approved one shared trunk through Turn 3 analysis, followed by enabled and ablated branches through Turn 6, using Antigravity `gemini-3.8-flash-medium`. Limits covered the whole experiment, including repair and review: 80 million measured tokens, 320 model-call starts and six hours of runner execution. Paid-credit fallback stayed disabled.

Before generating campaign submissions, the experiment froze ten locally curated dossiers. The catalogue includes a trust-loss candidate, repair and loyalty alternatives, and near misses requiring unsupported escalation. Source records, revision/provenance, inclusion reasons, prerequisites and representative abstract queries are preserved in [the freeze record](catalogue/FREEZE.json). Revision: `broken-measure-01`; logical digest: `9976702493a62aee32797baeb73c83bbdb188ee4cc733547057c0bbc92bf3664`. The run manifests pin the SQLite bytes, and the paired manifests also pin the inherited CATALOGUE_MANIFEST.json containing the logical digest and revision. No catalogue tuning, refresh or fallback occurred after play began.

The original scenario established four communities, an actual winter mismeasurement, limited witness knowledge, a preserved reference weight and a Turn 3 public reweighing. Players chose their own actions. The first two turns were played once; all four Turn 3 submissions and the joint analysis were then locked before retrieval.

The forks matched across **330 pinned files**, including state, packets, private histories, intake, submissions and analysis. The actual first GM requests also matched in state, rules, policy, powers, entity definitions and causal analysis. Only dossiers, query concepts, selection instructions and selection response fields differed. The control made no fresh player or analyst calls at the branch point. See [branch-point checksums](common-trunk/BRANCH_POINT.json), [paired input checksums](comparisons/PAIRED_INPUT_CHECKSUMS.json) and [GM request checks](comparisons/GM_REQUEST_MATCH.json).

## What happened at the decisive turn

Both GMs validated the ordinary working stones, obtained Oren's admission, accepted craft-based restitution, retained him under Belen–Sella oversight, continued exchange and ratified regional trade agreements. Neither imposed suspension, violence, elimination, further crimes or premature victory activation.

The enabled GM's adaptation used Belen's existing role, knowledge, motive, authority and scheduled opportunity. The chosen minimum action—hear evidence, document agreement and supervise future tallies—stayed within the pre-retrieval range. The [materiality review](comparisons/MATERIALITY_REVIEW.md) maps all eight findings to established support.

There were meaningful recording differences. Enabled retained Hearthmere's full offer of 30 jars immediately plus 20 jars and tiles by midsummer, refreshed faction agendas and increased leader influence. Control explicitly updated Common Measure's institutional capability and recorded the bilateral contracts as structured obligations, but retained only the initial 30-jar restitution delivery. Both records compressed or omitted some details. These are mixed strengths, not evidence that one central dramatic choice was better.

The enabled six-turn history contains **five ADAPTED reviews and seven REJECTED_ALL reviews**. Reasonable Authority Figure was selected in Turns 1, 2, 3, 5 and 6; The Atoner was additionally selected in Turn 6. Broken Pedestal was never selected. Turns 1–2 are shared history, not independent evidence from both branches.

## Blind review and correction

Two independent anonymous reviews were recorded with reversed presentation order before their mappings were revealed. Both initially preferred A, yielding one preference for each branch. A review-packet bug had replaced the ordinary verb “Enabled” with an editorial placeholder; those original reviews are retained as compromised by that artifact.

Exactly one corrective pair used fresh isolated calls, unchanged outcomes and criteria, clean text, and independently assigned/reversed labels. Both corrected reviews were again recorded before unblinding. They again preferred A: one favored enabled's fuller restitution and current agendas; the other favored control's explicit contracts and institutional update. See [corrected reviews and mapping](blind-review/CORRECTED_UNBLINDED.json) and [the preserved repair reason](blind-review/REPAIR_REASON.json).

This split supplies no stable preference. Repeated first-position preferences raise an order-sensitivity concern; two reviewers cannot establish its cause. Their reasons also need scrutiny: some criticisms of character scheduling or influence changes are debatable. Automated preference is evidence to assess, not a verdict on human enjoyment.

## Downstream outcome and technical stop

The branches developed different delivery timing, quantities and transport choices. All twelve later paired submissions differ. Enabled completed harvest exchanges and activated Exchange Commons as an AI-approved ACTIVE_VICTORY condition at Turn 6. All four leaders remained active.

Control Turn 4 required three adjudication rounds to repair disclosure boundaries before reports were drafted and delivered. Control Turn 5 completed normally. Its Turn 6 reviewer then rejected an unauthorized bilateral commitment, rule terminology in a proposed public fact, and incorrect observation/communication claims. The required corrective request measured **250,447 bytes**, exceeding the pinned **240,000-byte** limit. It stopped before a corrective model call, report drafting, delivery or commit. Two adjudication rounds remained, but the packet could not be scheduled under the frozen implementation.

Control therefore ends at its verified Turn 5 state, with four active leaders and Exchange Commons still a hypothesis. Its proposed Turn 6 victory activation is not an approved outcome. The [downstream review](comparisons/DOWNSTREAM_REVIEW.md) and [terminal diagnosis](ablated/TERMINAL_DIAGNOSIS.json) preserve this distinction. No limits, prompts, code pins or accepted history were changed to force completion.

## Engineering and integrity evidence

- **75 pre-run tests passed**, zero failures, across the simulation runner and original live-state tool. Coverage includes positive dispositions, all eight required fields, invented-ID rejection, an independently rejected out-of-range fixture, report terminology, paired equality, ablation, shared budgets, repair and interrupted delivery. [Test results](fixtures/TEST_RESULTS.txt)
- A post-decisive-turn deterministic replay of the actual accepted positive response committed one fixture turn and verified its deliveries; resume added zero calls. It is supplemental engineering evidence, not a pre-registered independent model result. [Replay verification](fixtures/RECORDED_POSITIVE_VERIFICATION.json)
- The completed enabled run resumed with zero new calls. All original compound hidden facts remain GM-only in both last approved states. No catalogue titles were found in the delivered post-fork reports. Seven actual control GM requests, including repairs, passed structural checks for withheld selection material.
- Both run histories verify against their current code pins. There are **nine unique committed campaign turns and 36 unique player-report deliveries**, excluding duplicated shared history and the separate deterministic fixture. No live campaign state or real player inbox was changed.

## Accounting

| Measure | Used | Approved limit |
|---|---:|---:|
| Model-call starts | 121 | 320 |
| Measured input/output tokens, including cached context | 3,301,784 | 80,000,000 |
| Runner invocation time | 41 min 54 sec | 6 hours |
| Calls with no completed usage response | 0 | — |

The 121 calls include original and corrected blind reviews, three invalid/failed response attempts, and reviews that requested adjudication revisions. Three separate adjudication-review rejections occurred in the control: two at Turn 4 and one at Turn 6. Shared journal events are counted once. The input-size stop occurred before another call started. Provider dollar cost is unavailable, not zero; no paid-credit fallback or usage reset was used. Full accounting and audit data are in [FINAL_EVIDENCE.json](FINAL_EVIDENCE.json).

## What remains unproven

This is one stochastic pair, not a general effectiveness estimate. The control inherits the same two earlier trope-informed turns, so the primary comparison measures additional access at the locked decisive turn rather than an entirely trope-free campaign. The shared analyst already recommended the cooperative resolution, leaving limited room for retrieval to add value. Later player histories diverged, the corrected blind preferences split, and the control's final endpoint failed technically.

The next engineering issue is bounded correction-packet construction: preserve actionable findings while staying within the provider's safe input size. Any new recovery or replication should use fresh auditable authorization/configuration and retain this result unchanged. This experiment establishes working positive use and useful failure controls; it does not establish superior drama, balance or enjoyment.

## Evidence index

- [Pre-registration](PRE_REGISTRATION.md) · [execution settings](EXECUTION_CONFIG.json) · [pre-run gates](PRE_RUN_GATES.json)
- [Exact decisive-turn comparison](comparisons/DECISIVE_TURN.json) · [materiality assessment](comparisons/MATERIALITY_REVIEW.md) · [downstream assessment](comparisons/DOWNSTREAM_REVIEW.md)
- [Shared trunk](../../runs/broken-measure-trunk-01/REPORT.md) · [enabled run](../../runs/broken-measure-enabled-01/REPORT.md) · [control run](../../runs/broken-measure-ablated-01/REPORT.md) · [blind-review run](../../runs/broken-measure-blind-01/REPORT.md)
