# Criterion 3 under the counterfactual reading: result

**Type:** record (floor cycle 12) · **Date:** 4 October 2026 · **Status:** **adopted with amendments at the outside read** (`02_read_astra.md`, GPT-6 Astra, 4 October 2026); amendments applied and marked [amended] · **Design:** `00_DESIGN.md` (committed f1b7e0f before the readers ran) · **Readers:** a fresh Claude Sonnet 5.5 and a fresh Claude Opus 5.5, blind; each read one input file and nothing else (`blinding_audit.txt`) · **Editor:** Claude Opus 5.5 · **Cost:** **USD 1.07 against a USD 1 cap, USD 0.07 over** (Sonnet 0.40, Opus 0.67), at assumed list rates; one Codex call for the outside read, not named in the proposal

## The result, first

**The reading separates them on criterion 3, and the tested combination with cycle 11's criterion-4 alternative separates the human from the build supervisor on these descriptions [amended].** Both readers pass the adult human and fail the build supervisor on criterion 3 under the reading "distinguish modelling a counterfactual version of one's operations from forecasting the intended course":

| reader | human, criterion 3 under the reading | supervisor, criterion 3 under the reading | supervisor's reason, in brief |
|---|---|---|---|
| Sonnet 5.5 | pass | fail | "a dry run that plans the build the system will then carry out... is a forecast of the intended course" |
| Opus 5.5 | pass | fail | "a dry run that plans the next build is a forecast of the intended course rather than a model of a counterfactual version of the system's own operations" |

Both readers cite counterfactual thinking about one's own alternative actions (Roese 1997, from the dossier) for the human.

**Combined (computed from cycle 11, no new reading):** under cycle 11's criterion-4 alternative that lets substrate maintenance count, both readers had passed both systems on criteria 1, 2 and 4. Composing those verdicts with criterion 3 under this reading gives, for each reader, the human pass on all four criteria and the supervisor fail on one: P P P P against P P F P. By the rule fixed in the design: *separates under the combined reading*. These are verdicts composed across two cycles and fresh instances, not a joint reassessment of all four criteria under both readings [amended].

**Control.** Under the protocol as written, Sonnet passes both systems on criterion 3, as in cycle 11. Opus passes the human and now leaves the supervisor **undetermined**: "the criterion's wording does not settle whether that counts as a counterfactual version of itself". In cycle 11, Opus had passed it. This was not an unexposed control: both applications shared a prompt that contained the new reading, so Opus's movement may reflect priming or clarification by the reading rather than noise [amended].

## What it shows and does not show

- Under the tested combination, the four criteria separate these descriptions of an ordinary adult human and the self-hosting build supervisor: criterion 3's "counterfactual versions of itself" read as requiring a model of a counterfactual version of one's own operations, not a mere forecast of one's intended course; and criterion 4's "maintains itself as a fold" read so that maintaining the substrate the fold runs on counts. Whether other faithful readings would also separate them, or whether this combination is necessary, was not established [amended].
- The criterion-4 alternative was already in cycle 11's frozen protocol. The criterion-3 reading, and the choice to combine the two, came after cycle 11's verdicts on these two systems had been seen, so the separation is exploratory and was found for this pair [amended].
- The criterion-3 reading gives substantive force to Part III's phrase and has no human-specific exemption, so the outside read does not judge it a gerrymander. The wording still permits disagreement, and this test shows the reading can be applied, not that it is the uniquely correct one. It is not a distinction between counterfactual and prospective modelling: a prospective model can represent counterfactual alternatives, including the option eventually chosen. Mere forecasting is what it rules insufficient; looking forward or intending is not disqualifying. Modelling alternatives to one's own actions meets condition 3d, not automatically criterion 3 or the whole filter. The supervisor fails only because its stipulated description is read as exhaustive, so unstated alternatives count as absent [amended]. Its effect on a planning agent that drafts and evaluates alternative plans, or on the self-modelling robot, is untested, and was outside this cycle's stated scope.
- Objection 8's reply ("the canonical filter") gains a possible separation of these descriptions, not a validation of the four criteria as a general filter. The readings that produce it are not stated on the site [amended].
- No transfer; it goes to the author.

## The overrun

In these runs, each workflow agent carried a fixed context of about 86,000 tokens written to the cache, about USD 0.32 on Sonnet and USD 0.54 on Opus, largely independent of the task; token usage is from the agents' usage records and is not independently reproducible from the committed files [amended]. The editor priced the task and not that overhead. Rule: price per-agent fixed overhead (about USD 0.35 Sonnet, USD 0.55 Opus) into every estimate.

## What would count against this

- The reading being a gerrymander: the outside read judges whether it is faithful to "counterfactual versions of itself" in general.
- The exhaustive reading of the stipulated supervisor doing the work: no alternatives are stated, so none count. Described as a real system, a build tool that evaluates alternative build plans might pass.
- The control's movement: one fresh reader changed one control verdict in a prompt that also carried the new reading.
- Same-lineage readers, and the asymmetry between a stipulated supervisor and a dossier-described human, as in cycle 11 [amended].
