# H1 across sampled codings, and where a paper-priority reading lands: results of steps A and B

**Type:** record of an exploratory analysis (lab type) · **Date:** 4 October 2026 · **Status:** **adopted with amendments at the outside read** (`95_read_astra.md`, GPT-6 Astra, 4 October 2026); amendments applied and marked [amended] · **Filing:** `80_ALL_CODINGS_FILING.md`. Committed order: filing 1b12833, script a9866f4 (after it reproduced two known results), step A results 33a37a1, step B input c6fdc3a, readers' codes and placement script ed485a1, placement results 81f90a1; git shows the order of committed files, not execution times [amended] · **Authors and models:** this record by Claude Opus 5.5 (editor); step B codes by a fresh Claude Opus 5.5 instance and a fresh Claude Fable 5.1 instance, blind · **Cost:** step A no agent spend; step B USD 3.08 at assumed list rates (Opus 5.5 0.71, Fable 5.1 2.37), of a USD 4 cap set by the author ("yes, go ahead with USD 4 for Opus and Fable", 4 October 2026); one Codex call for the outside read

## The result, first

**Reading-dependent, by the band fixed in the filing; no sampled coding met the reversal criterion [amended].** The analysis sampled 200,000 combinations of the codes that earlier coders had actually given to the disputed questions, in each of two universes [amended]. The share in which the pre-registered rule returns "Supported" (Zen minus TM at least 5 points, 95 percent interval above zero) is **19 percent in U1** (all four codings' candidate codes) and **27 percent in U2** (run 2's codings, from the brief that followed the pre-registration). That falls in the band "5 to 50 percent: reading-dependent".

| universe | Supported | Weakly supported | Inconclusive | Negative result | Counts against | sampled Δ: min, 2.5th, median, 97.5th, max |
|---|---|---|---|---|---|---|
| U1 (29 disputed items; 4 codings) | 19.1% | 53.5% | 23.5% | 3.9% | 0% | −0.09, 1.96, 4.06, 6.12, 8.13 |
| U2 (27 disputed items; run 2) | 26.7% | 56.5% | 16.1% | 0.7% | 0% | 0.60, 2.53, 4.39, 6.24, 8.13 |

In plain terms: the difference is in the predicted direction in all but 6 of the 200,000 sampled U1 codings and in all sampled U2 codings, with a typical size of about 4 points [amended]. Its 95 percent interval clears zero in 73 percent of U1 codings and 83 percent of U2 codings. It reaches the 5 points the test treated as mattering in about a fifth to a quarter of them. The minimum and maximum are those of the sample, not bounds on the universes [amended].

**Both universes keep known problem codes [amended].** They hold every code some coder gave, including item 76 as awareness-oriented, which the outside read found to reverse that item's meaning. A brief that followed the coding rule (run 2's) does not guarantee that every code it produced did.

**Run 3's coding was at the edge.** Its Δ of 1.23 lies below U1's 2.5th percentile (1.96). Its "Negative result" label is the outcome of 3.9 percent of U1 codings and 0.7 percent of U2 codings.

**Where a paper-priority reading lands (step B).** Two blind readers, Claude Opus 5.5 and Claude Fable 5.1, recoding from the paper's own account, agree on 25 of the 29 disputed questions. Both codings give Δ = 3.48 (95 percent intervals −0.34 to 7.30 and −0.15 to 7.11): **Inconclusive**, at the 30th percentile of U1.

## Step B: a paper-priority exploratory recoding [amended]

Two fresh readers, Claude Opus 5.5 and Claude Fable 5.1, each coded the 29 disputed questions. Each received only: the paper's own account of the two orientations (Part III, verbatim, no tradition named); the pre-registered definitions and coding rule; and the item wordings (`run3_prompts/paper_reader_input.md`). The audit (`stepB_blinding_audit.txt`) shows each made one read of its input file and returned structured output. Neither was told the hypothesis, the groups, the data, or that the items had been disputed. Every code either gave is one some earlier coder had given, so both codings lie inside U1.

**An unfiled instruction, the editor's [amended].** The input told the readers: "Where the definitions and the paper pull apart, follow the paper's meaning." The filing did not contain this precedence rule. It licenses departing from the registered constructs where the paper reads differently, so step B is a paper-priority exploratory recoding, not a demonstration that the registered meanings were preserved. Its effect on particular codes is untested.

**They agree on 25 of the 29.** They split on four: the three non-visual luminosity questions (brightness, clarity, radiance; Opus neither, Fable awareness) and the present moment (Opus neither, Fable content). Both code insight and item 76 as neither, the readings the outside read of run 3 asked for. Opus's reason for item 76, that a high score "marks the absence of the inward reflexive orientation", goes further than the wording: being different from reflexive awareness does not establish that inward orientation is wholly absent [amended].

**Codes the outside read flags [amended].** These concern which construct an item belongs to, not a reversal of what a high score means:
- Opus's present moment as neither: temporal self-location meets the registered criterion of time as content, and mentioning oneself does not make it oriented toward the modelling.
- Fable's three luminosity items as awareness-oriented: the wording does not settle that non-visual brightness, clarity or radiance indicate awareness knowing itself, and the registered tie-break favours neither.

| coding | content / awareness items | Δ | 95 percent interval | 90 percent interval | label | U1 percentile |
|---|---|---|---|---|---|---|
| Opus 5.5 reader | 22 / 13 | 3.48 | −0.34 to 7.30 | 0.28 to 6.68 | Inconclusive | 29.8th |
| Fable 5.1 reader | 23 / 16 | 3.48 | −0.15 to 7.11 | 0.43 to 6.52 | Inconclusive | 29.6th |
| the 16 combinations of their four splits | | 3.23 to 3.74 | | | 14 Inconclusive, 2 Weakly supported | |

**So the paper-priority reading lands below the distribution's median:** a difference of about 3.5 points in the predicted direction. Its 95 percent interval just reaches zero, and its upper end (about 7) leaves a difference of 5 or more open. By the pre-registered labels that is Inconclusive: neither a demonstrated difference nor a demonstrated absence of one that matters. The readers' four splits move Δ by about 0.51 points [amended]. **The label is fragile [amended]:** changing only Fable's item 66 from awareness to neither gives Δ = 3.65, 95 percent interval 0.011 to 7.29, which is "Weakly supported".

## Which questions it turns on

The table gives each disputed item's conditional mean Δ with the item coded each way, averaged over the other sampled choices. Those averages are not the largest effect a single item can have, and they do not test whether a label is stable: one item can tip a coding near a threshold, as item 66 does above [amended]. The largest conditional differences in U1:

| item (short label) | codes and mean Δ | difference |
|---|---|---|
| 92, breath stopping | content 3.56 · neither 4.55 | 0.99 |
| 78, resembling touch | awareness 3.51 · content 4.48 · neither 4.16 | 0.97 |
| 75, a quality of insight | awareness 3.70 · neither 4.41 | 0.71 |
| 28, a sense of self | awareness 4.40 · neither 3.71 | 0.69 |
| 86, the simplest kind of experience | awareness 4.34 · neither 3.76 | 0.58 |
| 29, part of inner life | awareness 4.31 · neither 3.80 | 0.51 |
| 2, alertness | awareness 3.83 · neither 4.27 | 0.44 |
| 76, different from awareness aware of itself | awareness 3.84 · neither 4.27 | 0.43 |

Two of the choices that lower Δ on average, insight as awareness-oriented and item 76 as awareness-oriented, are the two the outside read found to break the pre-registration's coding rule in run 3. The full table is in `results_all_codings.json`.

## What it shows and does not show

- It does not test H1. It is exploratory: the analysis was designed after results had been seen, though it chooses no coding.
- It shows that the earlier "Negative result" was not representative of the sampled codings. Most give a small difference in the predicted direction, and none met the reversal criterion. Whether the difference reaches the 5 points the pre-registration set depends on how the disputed questions are read, and near the threshold a single question can tip the label.
- A paper-priority reading by two blind readers lands at about 3.5 points, Inconclusive, below the distribution's median and above run 3's coding.
- Uniform sampling treats every combination of readings as equally reasonable, which they are not. The choice of universes and their weighting matters to the shares [amended].
- The pre-registration's limits stand: the questionnaire asks about the awareness end, TM is a proxy for Advaita, participants self-selected, and a difference in these reports says nothing about the fold or whether there are exactly two orientations.

## What would count against this

- The universes hold only codes some coder gave. A reading no coder proposed is not represented, and known problem codes are kept.
- The 63 items fixed by agreement were inherited; the step B readers did not reassess them [amended].
- Two readers agreeing does not establish construct validity or a uniquely correct reading of the paper [amended].
- The variation across codings is specification variation, not sampling uncertainty; the shares are not probabilities that H1 is true. A share under 5 percent would not have meant that no fair reading supports a difference [amended].
- The editor designed this after seeing that run 3's result was coding-sensitive. The filing fixed the universes, the sampling and the reading bands before the script was committed, and the script reproduced two known results before it sampled.
