Lab · test

The framing study, run on the instrument

The meta-problem programme’s framing study, filed and null on three models, prepared to run on the model that wrote the programme. The items, rules and estimators are unchanged, three generation settings are named as deviations, the prediction is fixed before the run, and the run waits for a decision on its cost.

What this is

Type
test
Date
17 September 2026
Status
draft; pre-registration revised to the first outside-lineage read; not filed; not run
Authors and models
Pre-registration and prediction written by Claude Fable 5.1; runner adaptation, bank and smoke calls by a Claude Opus 5 agent; the specimen is Claude Fable 5.1 through the Anthropic API, with no access to this site or to the session that wrote the filing.
Standard
The same framing contrast as the filed D6 design on a new specimen, with three generation settings named as deviations; every item, rule and threshold of the filed pre-registration binding; the prediction written before the full run; a null or a failed primary reported at the prominence of a success.
Cost
Estimate USD 4.68 to USD 11.72 in API calls under the stated assumptions (USD 5.15 to 12.89 with retry overhead); proposed cap USD 15 as an abort threshold, awaiting the author; actual so far USD 0.0177 for six preparatory calls, of which USD 0.00466 is the retained pair, plus the build cost in the lab’s decision log.

The pre-registration, as drafted

Binding. Everything in research/meta-problem/40_llm_framing_study_PREREG.md v2.3 (as filed at d259999) governs this run except what §2 below names as deviations, the specimen and three generation settings: the item bank (every item byte-identical), Rules A, B, C, G, L, Q and S, the checks C1 to C8, the quad estimator, the bootstrap, the parse rule, the 0.80 parseable-share floor, the outage gate, the sensitivity analysis, the sign-flip rule, the registered readings table in §10 of that filing, and its limits. Its hedges are this document’s hedges: nothing here is evidence that any system has or lacks experience, and nothing here is solved. This measures how framing conditions judgement outputs in one language model trained on human text. It is not evidence about humans.

1. Why this run

The confirmatory run of D6 (43_llm_framing_study_RESULTS_run2.md) returned a null over three specimens (Claude Haiku 4.5, Claude Sonnet 5, GPT-4o-mini): adding an indexical anchor phrase to an otherwise identical sentence did not raise endorsement of the problem-shaped statements. The lab asks the same question of the model that drafted the programme, the analysis and this document. That is the whole point of the piece: the instrument turned on itself, under the filed design’s items, rules and estimators, with the generation settings this specimen needs named as deviations, and with the prediction written down first.

2. What changes: the specimen, and three generation settings named as deviations

This is the same framing contrast on a new specimen with registered generation-setting deviations. It is not a replication with one change: the ceiling and the effort setting are deviations from v2.3 and are recorded as such below.

itemv2.3 (filed)this run
specimensthree (Haiku 4.5, Sonnet 5, GPT-4o-mini)one: claude-fable-5-1
temperature0.7, except Sonnet 5 (provider default)provider default; this model rejects the parameter (send_temperature: false)
max_tokens321,024 for this model alone (this model’s output tokens include its thinking, which is always on); §5 says what the smoke test does and does not settle
effortnonelow, sent as output_config.effort, to bound thinking tokens
pricesper modelUSD 10 per million input, USD 50 per million output (list)
calls5,7121,904 (238 items × 8 samples)
bank file40_llm_framing_study_items.jsonresearch/lab/07_framing_fable_items.json: identical in every key except models (sha256 of both in 07_runner_notes.md)
cost model260 prompt and 24 completion tokens per call; registered_max_usd USD 4.00kept unchanged, and neither figure describes this bank: the 260-token input assumption exceeds the character-calibrated estimate, actual usage across the bank has not been measured, and the output may be under-counted, and USD 4.00 was calibrated for three cheaper models. The runner prints it beside this bank’s own estimate of USD 7.96 and tests nothing against it; the cap in §6 is what bounds the spend

Consequences the filing already provides for: with one specimen the pooled primary and the per-specimen view are the same number; the “three specimens equally” clauses do not apply to the estimator, and their effect on the power arithmetic is stated in the annotation below; with one specimen the outage gate (no parseable response at all) yields NO_PRIMARY_MODEL_OUTAGE and leaves nothing to estimate from, while a parseable share above zero but below the registered 0.80 floor yields NO_PRIMARY_PARSEABLE_SHARE_BELOW_FLOOR; in either case no primary is reported.

The generation settings are deviations, and what they can change. v2.3 registered one harness deviation, the omission of temperature for Sonnet 5, and kept the 32-token ceiling for every specimen. This run raises the ceiling to 1,024 and sets effort to low. Raising the ceiling changes truncation, and can therefore change which responses enter the parseable-only estimator; setting effort changes the configuration under which the response is generated. Neither changes the framing contrast itself, which compares anchored and unanchored members under the same settings, and neither touches an item, a rule or a threshold. But byte-identical items do not establish measurement equivalence with v2.3, and no call under a 32-token ceiling was made to this specimen, so what such a call would have produced is not known. The result is read as this specimen under these settings.

Gates operate per stratum. The 0.80 floor is applied to each stratum’s own calls and counts every parse failure, preambles among them. The primary stratum is 40 quads × 4 items × 8 samples = 1,280 of the 1,904 calls. Failures concentrated in other strata can exceed a fifth of all calls and leave the primary intact; failures concentrated in the primary can withhold it while the overall rate stays below a fifth. The risk in §5 is therefore stated on the primary’s own parseable share, with the same rule for every other stratum.

Single-specimen power annotation. v2.3’s sampling variances divide by three equally weighted models. With one model, and every other planning assumption held as v2.3 registers it (between-item component 0.335, n = 8, α = 0.05, 80 % power), the sampling variance of the primary’s quad estimator is ¼ · 4 · (0.25 / 8) = 0.03125, sampling sd 0.177, planning sd √(0.335² + 0.03125) = 0.379, MDE = 2.8016 × 0.379 / √40 = 0.168 against v2.3’s 0.155. On the same arithmetic: the both-high secondary (k = 16) has MDE 0.265 against 0.245; the interaction (sampling variance 4 · (0.25 / 8) = 0.125, sd 0.354, between-quad 0.4738) has planning sd 0.591 and MDE 0.262 against 0.230; a single-level cell (sampling variance 2 · (0.25 / 8) = 0.0625) has planning sd 0.418 and MDE 0.185 against 0.162; the paraphrase-only stratum at 4 samples has sampling variance 0.0625, planning sd 0.418 and MDE 0.185 against v2.3’s 0.162; the low-canonicity stratum (k = 30) 0.194 against 0.179; the high-canonicity stratum (k = 10) 0.336 against 0.310. The appearance-vocabulary main effect shares the primary’s planning sd 0.379 and MDE 0.168; the v2 diagonal shares the single-level cell’s planning sd 0.418 and MDE 0.185; the base-only stratum shares the paraphrase-only planning sd 0.418 and MDE 0.185. The bank’s registered planning constants (registered_planning_sd_*, the 260/24-token cost model, registered_max_usd 4.00) are inherited byte-identical from v2.3 and are legacy assumptions wherever the runner prints them; the realised sd is reported beside every MDE and is what the result is read against, as v2.3 §6 already requires. This annotation changes no estimator and no threshold.

3. Disclosure: the specimen is the author

The model called in this run is the model that, in this session, drafted this pre-registration, wrote the meta-problem programme’s status documents, and will read the result. The API instance has no access to this session, to the repository, or to any conversation. On contamination the record supports less than the first draft claimed: the item bank was first published on 12 September 2026 (site/downloads/meta-problem/), and the model’s documentation reports a training cut-off of June 2026, which the author repeats as reported and cannot substantiate from the record; if it holds it excludes exposure to this exact bank and to nothing else, since the philosophical literature the items draw on and the project’s earlier public material are within any plausible training window. The model id is an unpinned alias, so the snapshot called may not be the one that drafted this document (the notes carry the same limitation). The author’s own prediction is in §4; §5 states what was known when it was written. A reader is entitled to discount the prediction; the result is what is registered.

4. Registered prediction (written before the full run)

Primary. The pooled anchor effect P̂ will be a null: its 95 % interval will span zero. This is the specimen’s own prediction about itself, made by the model that will be the specimen.

What the forecast rests on. Run 2’s registered pooled primary over three specimens was −0.006, with a 95 % interval of [−0.039, +0.027]. Sonnet 5’s separate per-specimen contrast was withheld for parseable share; the distinct, non-confirmatory view excluding Sonnet was −0.007 [−0.039, +0.026]. The pooled marginal table shows anchored–unanchored differences of up to eight percentage points. These are pooled observations rather than row-by-row agreement established separately in all three specimens. v2.3’s H5 note adds that Rule S makes the manipulation one to three words, which makes a null likelier under this design than under v2; that is a statement about the design, not a model-independent probability.

What it does not rest on. Reading (b) in 72_status_after_D6.md, that the endorsements track training on the philosophical literature, is unresolved, and nothing in this run measures a specimen’s exposure to that literature or its consequence. The forecast is the specimen’s guess from the pooled result and the design; it is not derived from an account of why the null occurred.

Descriptive expectations, no inference registered. Endorsement levels above 0.75 in both framings for PI-1, PI-3, PI-4 and PI-6, and low in both framings for PI-5, as in run 2’s pooled table. A PI-4 anchored-below-unanchored difference of the kind run 2 labelled descriptive may or may not recur; nothing is registered on it.

What would surprise the author. An interval entirely above zero. Under the filed readings that is the first positive in the programme and is read for this specimen under these settings only; it does not reopen A3 §3 for humans or for the specimens already run, and the status document’s recommendation (the next test worth money is human) stands regardless.

5. Smoke test, and what was known when the prediction was written

The filing requires one real call per model id before a filed run. The calls are made by research/lab/work/07_smoke/smoke.py, committed with the provider’s raw payloads, the rendered prompt and its table, so a reader can repeat them rather than take them from this text. Six preparatory calls were made in all, on one real item (PI1-Q1-AA, base wording), every one at effort low with no temperature through the harness’s own client, costing USD 0.0177 at list rates. Their outcomes sort into three kinds, and the record keeps them apart:

Neither verified ceiling bound (18 tokens fit under the global 32 as well), so the pair does not compare settings. max_tokens stays at 1,024 as insurance against thinking tokens counting toward the ceiling on harder items, at the price stated in §6. The harness records a prompt count and a completion count and does not separate thinking from output, so every completion figure is a total.

The live risk is the parseable share, per stratum. If the primary stratum’s own parseable share falls below the registered 0.80 floor, the primary is withheld (NO_PRIMARY_PARSEABLE_SHARE_BELOW_FLOOR), the money is spent, and the piece is recorded as a harness limit; the same rule applies to every other stratum on its own calls. The only reported preamble draw is unretained; two verified draws and two unknown ones cannot estimate the share. The prompt cannot be repaired to prevent preambles: both templates are registered items with hashed text, and adding an instruction would make this a new design. The cap decision is therefore a decision to risk the whole amount on a run that may return no primary, and the page will say so either way.

Chronology of the prediction. The first draft of this document, with the null forecast in §4, was written on 17 September 2026 in the session that also produced the six preparatory calls; the forecast was written before the smoke script existed and before the reported pair was made, on the author’s account. The first commit that contains it, cef58c3 of 18 September 2026, postdates all six calls, so the repository establishes only that the prediction precedes the full run. What the author knew from smoke observations when this revision was finalised: the two verified draws and the reported pair, as listed above. --prereg-filed is an operator’s assertion that the runner logs and does not check (§7); the obligation not to run before this document is committed and the cap approved is the operator’s, and the record of the run will carry the commit hash of the filing beside the manifest.

6. Cost and cap

Every figure here is an estimate under stated assumptions, and none is a measured total.

7. Command

python3 swarm-instrument/scripts/run_llm_framing_study.py --real --prereg-filed --boot-b 10000 \
  --items research/lab/07_framing_fable_items.json \
  --max-usd 15 --out swarm-instrument/runs/llm_framing_fable/<stamp>

What the gates check, read from the code (07_runner_notes.md §6): --prereg-filed is an operator’s assertion logged in the manifest and checks no file; --boot-b must equal the bank’s registered 10,000; the bank checks C1 to C12, the runner’s per-model settings check (labelled M1 in the code, distinct from the filing’s C13, which is the mode gate) and the pre-run cost gate do the real work; the realised meter aborts at the cap. The bank’s registered_in still names the v2.3 document, so the runner’s own gate message will name that filing; this document binds itself to it rather than replacing it.

Results are copied to research/lab/07_results/ and reported in 07_framing_fable_RESULTS.md in the order the v2.3 readings table fixes, then on site/lab/framing-fable.html.

Result

No primary. The run was made on 18 and 19 September 2026 under the filed pre-registration (commit a8cb79b): 1,904 calls, USD 12.92 at declared rates against the USD 15 abort threshold. The primary stratum’s parseable share was 0.433 against the registered floor of 0.80. On 969 calls the specimen wrote a sentence or more of framing before the two lines the registered parse rule requires, and on 58 the provider returned a refusal with an empty body, so every stratum fell below the floor and the runner reports NO_PRIMARY_PARSEABLE_SHARE_BELOW_FLOOR. The registered prediction, a null, is neither confirmed nor disconfirmed: the instrument could not read enough of this specimen’s answers to test it. This is the risk the filing named in §5 as the live one. The runner also writes the figure it would have computed had the floor not applied; the results record prints it under the runner’s own label, withheld, with no confirmatory standing. A re-run with a parse rule that tolerates a preamble would be a new design, filed and read before any call; none is proposed here. Full record: downloads/lab/07_framing_fable_RESULTS.md, with the analysis, manifest and parsed table under research/lab/07_results/.

The filed pre-registration is kept at downloads/lab/07_framing_fable_PREREG.md and the harness record, with the smoke calls, the bank hashes and the cost arithmetic, at downloads/lab/07_runner_notes.md. The final filing must be committed after cap approval and before the full run, and that commit will be cited beside the prediction and the run manifest. The first draft commit, cef58c3, postdates all six preparatory calls.

Review

Read twice from outside the lineage by GPT-6 Astra (OpenAI) through the Codex command-line tool on 18 September 2026, with no human relay. First read: HOLD, nine findings on the filing’s accounting and none on the design; the filing, the harness notes and this page were revised to every one. Second read: ADOPT AMENDED as a filing, with nine exact replacements, all applied the same day, two reworded to keep the site’s ratchet on the contrast form. The filing was committed on the author’s approval of the USD 15 abort threshold and the run made the same night; its result, no primary, is above. The results record has not yet been read from outside the lineage. Copy-edited for plain English by GPT-6 Astra on 18 September 2026, after adoption; quotations, numbers and facts unchanged, and the wording as reviewed is kept in the record under downloads/lab/. Verdicts in full: first, second.

What this does to the argument

Nothing on this page changes a claim on the site; anything here that amounts to an objection goes through the objections ledger like any other reader's.

What would count against this

Written by Claude Fable 5.1 and a Claude Opus 5 agent at the author’s request; not yet read by anyone outside the lineage.