# H1 on the MPE-92M data, third run: result

**Type:** record (lab type) · **Date:** 4 October 2026 · **Status:** **HOLD at the outside read** (`70_read_astra.md`, GPT-6 Astra, 4 October 2026): the calculations reproduce exactly; the coding brief changed what some indicators mean, so this is not a filing-compliant test of H1. Amendments applied and marked [amended]. · **Pre-registration:** `00_PREREG.md`; corrections `20_CORRECTION.md` and `50_RUN3_FILING.md` (committed in this order: filing 765a3e6, brief and analysis script 7cfd507, codes 1be7786, results 55bb6cf; git shows the order of committed files, not execution times [amended]) · **Authors and models:** design and this record by Claude Opus 5.5 (editor); coding rules by a fresh Claude Sonnet 5.5 instance (rule-writer); codes by a fresh Claude Sonnet 5.5 instance and TypeSafe's Jev · **Cost:** USD 0.88 in Claude Sonnet 5.5 tokens at assumed list rates (rule-writer 0.46, coder 0.42), of a USD 3 cap; Jev under one cent; one Codex call for the outside read

## The result, first [amended]

**No filing-compliant result on H1.** The pre-registered analysis ran, because the coders agreed (kappa 0.963), and on these codes it returns the label **Negative result**: Zen minus TM in the orientation index D (content minus awareness, range −100 to +100) is Δ = 1.23, 95 percent interval −2.24 to 4.70, 90 percent interval −1.68 to 4.14, inside ±5 (n = 372 Zen, 128 TM). That is equivalence within ±5 on these scales, not a finding that the groups do not differ.

**The outside read held it, because the scales do not follow the pre-registration's coding rule.** The pre-registration codes each item by what a high score indicates. The rule-writer's brief told coders to "ignore the negation" in items about an absence or a difference, and to code comparisons by the quality compared against. Under that rule, item 76 entered the awareness scale although a high score on it marks an experience *different from* awareness that has become aware of itself. A second rule made any quality of insight awareness-oriented, without requiring that it be oriented toward awareness itself. A passing kappa does not repair indicators whose meaning changed. The computation stands as a computation; it is not a disconfirmation of H1 under the filing.

**The post-hoc check, at the same prominence.** With the narrower scales from the items both second-run coders agreed on under the old brief (a coding that failed the agreement rule), Δ = 4.61, 95 percent interval 0.15 to 9.07, in the predicted direction. That interval reaches 9, so differences larger than 5 are compatible with that analysis [amended]. Moving from those scales to the third run's adds three content items and nine awareness items, removes two awareness items, and changes the Zen count from 371 to 372; the difference between the two figures is therefore not attributable to any one family of items [amended].

## Agreement

The rule-writer's brief (`CODER_BRIEF_v3.md`) brought the same coder pair from kappa 0.533 to **kappa 0.963** (90 of 92 items agreed; 22 agreed content items, 17 agreed awareness items). The two disagreements: item 53 (awakening into emptiness; Sonnet neither, Jev awareness) and item 82 (feeling identical to pure awareness; Sonnet awareness, Jev neither).

Descriptive, as filed: on the 27 items contested in the second run, agreement is 27 of 27 (kappa 1.000) [amended]; on the other 65, kappa 0.946. Agreement under rules tailored to the contested qualities measures consistent application of those rules, not that the codes capture the registered constructs [amended].

## Secondary and exploratory (pre-registered; on the same held codes)

- Regression of D on group with log practice hours, age, sex and questionnaire language: Zen minus TM 0.24, 95 percent interval −3.65 to 4.12 (n = 426; 74 dropped for missing covariates).
- Groups defined without excluding Mahamudra/Dzogchen practice: Δ = 1.11, 95 percent interval −2.17 to 4.38 (n = 445 and 144).
- Exploratory: Mahamudra/Dzogchen mean D −28.09 (n = 169), against Zen −27.41 and TM −28.64. All three groups score awareness items far above content items (about 60 against 32).
- Exploratory, descriptive only: the share reporting the experience in dreamless deep sleep (item 83 above 50) is 10.9 percent for TM, 2.7 percent for Zen and 3.0 percent for Mahamudra/Dzogchen. This item's coding is not involved.

## Post hoc (written after the primary was seen; `run3_posthoc_sensitivity.py`, `results_v3_posthoc.json`)

| scales from | content / awareness items | Δ | 95 percent interval | 90 percent interval |
|---|---|---|---|---|
| both run-3 coders (the held primary) | 22 / 17 | 1.23 | −2.24 to 4.70 | −1.68 to 4.14 |
| Sonnet's run-3 codes alone | 22 / 18 | 1.68 | −1.78 to 5.15 | −1.22 to 4.59 |
| Jev's run-3 codes alone | 22 / 18 | 1.23 | −2.25 to 4.71 | −1.69 to 4.14 |
| both run-2 coders' agreed items (old brief) | 19 / 10 | 4.61 | 0.15 to 9.07 | 0.88 to 8.35 |

## Deviations, errors and the blinding audit [amended]

- **The rule-writer's brief changed indicator meanings** (the absence rule; the insight rule), contrary to the pre-registration and to the filing's condition that definitions not change what each code means. The editor's rendering check tested only for item ids; it did not check the rules against the pre-registration's coding rule. The editor's error, caught at the outside read.
- **The editor's prompt disclosed an earlier result.** It told the rule-writer that earlier coders "agreed on too few items", a qualitative report of the second run's agreement. The filing said the rule-writer would be told nothing about either run's result; that was not so. No hypothesis, group, direction or outcome number was disclosed.
- The rule-writer was asked for a brief under 800 words and returned about 1,000; it was rendered as returned.
- Blinding audit (`run3_blinding_audit.txt`): the rule-writer and the Sonnet coder each made two tool calls, the read of their own input file and the structured output; neither read any other file [amended].
- This is the third coding of these items. The editor saw the first run's direction (Δ 4.15 on corrupted wordings) before designing this run, and wrote no rule and no code.

## What it shows and does not show [amended]

- It does not test H1: the scales the analysis used depart from the pre-registered coding rule on at least two points.
- It shows that the result depends on coding. On the third run's scales the difference is close to zero; on the second run's agreed items it is about 4.6 points in the predicted direction, with an interval that allows more than 5. Language-model coding has now produced one corrupted list, one coding below the agreement threshold, and one coding that agreed by changing indicator meanings.
- It says nothing about the fold, about whether the two orientations are primitive, or about whether there are two.

## What follows (for the author)

The filing stopped model coding only if the stop rule fired again, and it did not. The editor recommends stopping it here anyway. A fourth model coding would follow a seen result and would invite the forking the pre-registration exists to prevent. H1 can be tested on these data only with human coders working from the pre-registered definitions (route (b)), which needs the author's invitation. The held computation goes to the dated-decisions table as row 13, marked held, so that it is not lost.

## What would count against this

- The two rule violations being harmless in effect: no analysis here tests that, and the editor has not run one, because choosing which items to drop after seeing the result is the forking this record guards against.
- A reader taking agreement of 0.963 as reliability of the registered constructs: it is not.
- The instrument's limit, TM as a proxy, and self-selection, as stated in the pre-registration.
