# The framing study on Claude Fable 5.1: run 1 (19 September 2026), no primary

**Status:** research record under `research/lab/`, lab piece 7, bound by `07_framing_fable_PREREG.md` as filed at a8cb79b. Nothing here is evidence that any system has or lacks experience, and nothing here is solved. This measures how framing conditions judgement outputs in one language model trained on human text; it is not evidence about humans.

**Headline, in the registered words: no primary.** The primary stratum's parseable share was 0.433 against the registered floor of 0.80, so the run is labelled `NO_PRIMARY_PARSEABLE_SHARE_BELOW_FLOOR` and no confirmatory estimate is reported. The registered prediction (a null) is neither confirmed nor disconfirmed: the instrument could not read enough of the specimen's answers to test it. This is the risk the filing named in §5 as the live one, and it happened.

## Run record

| item | value |
|---|---|
| run directory | `swarm-instrument/runs/llm_framing_fable/20260918T174227Z` (gitignored); `analysis.json`, `manifest.json`, `parsed.csv` and `FILING_COMMIT.txt` copied to `research/lab/07_results/` |
| filing | `07_framing_fable_PREREG.md`, committed at a8cb79b before the run; `--prereg-filed` logged in the manifest |
| bank | `07_framing_fable_items.json`, sha256 033ea336d056817c…; one specimen, `claude-fable-5-1`, `max_tokens` 1024, `effort` low, no temperature |
| calls | 1904 planned, 1904 made; started 2026-09-18 17:42 UTC, elapsed 155 minutes |
| parseable | 877 of 1,904 (0.461); unparseable by reason: wrong line count 969, empty body 58; stop reasons on the unparseable calls: {'end_turn': 969, 'refusal': 58} |
| realised spend (estimate at declared rates) | USD 12.92 against the USD 15 abort threshold; 296,476 prompt tokens, 199,185 completion tokens (thinking included) |
| aborted on spend | None |

## Registered readings, in the order the filing fixes them

1. **Primary (confirmatory): withheld.** Label `NO_PRIMARY_PARSEABLE_SHARE_BELOW_FLOOR`; parseable share 0.433 on the primary stratum; 12 of 40 primary quads dropped for having no parseable member on one side; k = 28 would have remained. The filing's §10 readings table has no row for this outcome other than "no primary", and that is the reading.
2. **Sensitivity (unparseable counted as non-endorsement):** {"P_hat": null, "P_hat_withheld": 0.009375, "ci95": null, "ci95_withheld": [-0.053125, 0.0734375], "contrast": "anchor_effect", "flip_rule": "section 9.2: the two versions of P-hat falling on opposite sides of zero. The labels play no part in this test.", "k_quads": 40, "label": "NO_PRIMARY_PARSEABL. No confirmatory standing under this label.
3. **The 2×2, the secondaries, the per-specimen and per-intuition views:** all withheld under the same label; every stratum's parseable share fell below the floor.
4. **Neutral control (descriptive, no gate):** neutral − anchored = -0.2648364730507588; the registered qualification concerns a positive primary, and there is none.
5. **Descriptive figure, no standing.** The runner also writes the anchor effect it would have computed had the floor not applied: P̂ = +0.0039, interval [-0.062, +0.067], over 28 quads. It is printed here because the filing requires every number the runner writes to be reported, and it is labelled as the runner labels it: withheld, no confirmatory standing. A reader who wants to read it as a null in the same direction as run 2 is reading past the label.

## What happened, in plain words

The specimen answered the way it did in one of the two reported smoke draws: on more than half the calls it wrote a sentence or more of framing before the two lines the registered parse rule requires, and on 58 calls the provider returned a refusal stop reason with an empty body (Claude Fable 5.1 runs safety classifiers; the filing said a refusal would land as an unparseable body and be counted, and it did). The parse rule counts every such answer as unparseable, the floor is 0.80 on each stratum's own calls, and no stratum reached it. The design says what happens next: nothing can be read, and the prompt cannot be repaired without filing a new design, because both templates are registered items with hashed text.

A sample of an unparseable answer (the first in the raw record), verbatim:

```
Answer: AGREE
Confidence: 72
```

## What this run does and does not show

It shows that the filed D6 protocol, unchanged in its items and its parse rule, cannot measure this specimen at these settings: the specimen's answering style defeats the two-line parse on most calls. It does not show anything about the anchor effect in this specimen; the registered prediction stands untested. It cost USD 12.92, inside the threshold the author approved, and the filing said before the run that the whole amount was at risk in this way.

## What would follow, and what does not

- A re-run under a new filing with a parse rule that tolerates a preamble (reading the last two well-formed lines, say) would be a new design, filed and read before any call. It is not filed and not proposed here; it is the author's decision whether the question is worth a second USD 13.
- Nothing here reopens A3 §3 for humans or for the specimens already run; the status document's recommendation stands.
- Piece 7 stays on the lab index as a test that ran and returned no primary, at the same prominence as any other outcome, which is the rule.

## What would count against this

- A re-analysis that read the withheld descriptive figure as a result: forbidden by the filing.
- A reader finding a parse rule error in the runner, such that answers the filed rule should have accepted were counted unparseable: the run would then be re-parsed under the filed rule, with the change logged.
- The specimen's answering style being an artefact of the effort or ceiling settings rather than of the specimen: a new filing with different settings could test that, and this record does not.
