Lab · test
The framing study, run on the instrument
The meta-problem programme’s framing study, filed and null on three models, prepared to run on the model that wrote the programme. The items, rules and estimators are unchanged, three generation settings are named as deviations, the prediction is fixed before the run, and the run waits for a decision on its cost.
What this is #
The pre-registration, as drafted #
Binding. Everything in research/meta-problem/40_llm_framing_study_PREREG.md v2.3 (as filed at d259999) governs this run except what §2 below names as deviations, the specimen and three generation settings: the item bank (every item byte-identical), Rules A, B, C, G, L, Q and S, the checks C1 to C8, the quad estimator, the bootstrap, the parse rule, the 0.80 parseable-share floor, the outage gate, the sensitivity analysis, the sign-flip rule, the registered readings table in §10 of that filing, and its limits. Its hedges are this document’s hedges: nothing here is evidence that any system has or lacks experience, and nothing here is solved. This measures how framing conditions judgement outputs in one language model trained on human text. It is not evidence about humans.
1. Why this run #
The confirmatory run of D6 (43_llm_framing_study_RESULTS_run2.md) returned a null over three specimens (Claude Haiku 4.5, Claude Sonnet 5, GPT-4o-mini): adding an indexical anchor phrase to an otherwise identical sentence did not raise endorsement of the problem-shaped statements. The lab asks the same question of the model that drafted the programme, the analysis and this document. That is the whole point of the piece: the instrument turned on itself, under the filed design’s items, rules and estimators, with the generation settings this specimen needs named as deviations, and with the prediction written down first.
2. What changes: the specimen, and three generation settings named as deviations #
This is the same framing contrast on a new specimen with registered generation-setting deviations. It is not a replication with one change: the ceiling and the effort setting are deviations from v2.3 and are recorded as such below.
| item | v2.3 (filed) | this run |
|---|---|---|
| specimens | three (Haiku 4.5, Sonnet 5, GPT-4o-mini) | one: claude-fable-5-1 |
| temperature | 0.7, except Sonnet 5 (provider default) | provider default; this model rejects the parameter (send_temperature: false) |
| max_tokens | 32 | 1,024 for this model alone (this model’s output tokens include its thinking, which is always on); §5 says what the smoke test does and does not settle |
| effort | none | low, sent as output_config.effort, to bound thinking tokens |
| prices | per model | USD 10 per million input, USD 50 per million output (list) |
| calls | 5,712 | 1,904 (238 items × 8 samples) |
| bank file | 40_llm_framing_study_items.json | research/lab/07_framing_fable_items.json: identical in every key except models (sha256 of both in 07_runner_notes.md) |
| cost model | 260 prompt and 24 completion tokens per call; registered_max_usd USD 4.00 | kept unchanged, and neither figure describes this bank: the 260-token input assumption exceeds the character-calibrated estimate, actual usage across the bank has not been measured, and the output may be under-counted, and USD 4.00 was calibrated for three cheaper models. The runner prints it beside this bank’s own estimate of USD 7.96 and tests nothing against it; the cap in §6 is what bounds the spend |
Consequences the filing already provides for: with one specimen the pooled primary and the per-specimen view are the same number; the “three specimens equally” clauses do not apply to the estimator, and their effect on the power arithmetic is stated in the annotation below; with one specimen the outage gate (no parseable response at all) yields NO_PRIMARY_MODEL_OUTAGE and leaves nothing to estimate from, while a parseable share above zero but below the registered 0.80 floor yields NO_PRIMARY_PARSEABLE_SHARE_BELOW_FLOOR; in either case no primary is reported.
The generation settings are deviations, and what they can change. v2.3 registered one harness deviation, the omission of temperature for Sonnet 5, and kept the 32-token ceiling for every specimen. This run raises the ceiling to 1,024 and sets effort to low. Raising the ceiling changes truncation, and can therefore change which responses enter the parseable-only estimator; setting effort changes the configuration under which the response is generated. Neither changes the framing contrast itself, which compares anchored and unanchored members under the same settings, and neither touches an item, a rule or a threshold. But byte-identical items do not establish measurement equivalence with v2.3, and no call under a 32-token ceiling was made to this specimen, so what such a call would have produced is not known. The result is read as this specimen under these settings.
Gates operate per stratum. The 0.80 floor is applied to each stratum’s own calls and counts every parse failure, preambles among them. The primary stratum is 40 quads × 4 items × 8 samples = 1,280 of the 1,904 calls. Failures concentrated in other strata can exceed a fifth of all calls and leave the primary intact; failures concentrated in the primary can withhold it while the overall rate stays below a fifth. The risk in §5 is therefore stated on the primary’s own parseable share, with the same rule for every other stratum.
Single-specimen power annotation. v2.3’s sampling variances divide by three equally weighted models. With one model, and every other planning assumption held as v2.3 registers it (between-item component 0.335, n = 8, α = 0.05, 80 % power), the sampling variance of the primary’s quad estimator is ¼ · 4 · (0.25 / 8) = 0.03125, sampling sd 0.177, planning sd √(0.335² + 0.03125) = 0.379, MDE = 2.8016 × 0.379 / √40 = 0.168 against v2.3’s 0.155. On the same arithmetic: the both-high secondary (k = 16) has MDE 0.265 against 0.245; the interaction (sampling variance 4 · (0.25 / 8) = 0.125, sd 0.354, between-quad 0.4738) has planning sd 0.591 and MDE 0.262 against 0.230; a single-level cell (sampling variance 2 · (0.25 / 8) = 0.0625) has planning sd 0.418 and MDE 0.185 against 0.162; the paraphrase-only stratum at 4 samples has sampling variance 0.0625, planning sd 0.418 and MDE 0.185 against v2.3’s 0.162; the low-canonicity stratum (k = 30) 0.194 against 0.179; the high-canonicity stratum (k = 10) 0.336 against 0.310. The appearance-vocabulary main effect shares the primary’s planning sd 0.379 and MDE 0.168; the v2 diagonal shares the single-level cell’s planning sd 0.418 and MDE 0.185; the base-only stratum shares the paraphrase-only planning sd 0.418 and MDE 0.185. The bank’s registered planning constants (registered_planning_sd_*, the 260/24-token cost model, registered_max_usd 4.00) are inherited byte-identical from v2.3 and are legacy assumptions wherever the runner prints them; the realised sd is reported beside every MDE and is what the result is read against, as v2.3 §6 already requires. This annotation changes no estimator and no threshold.
3. Disclosure: the specimen is the author #
The model called in this run is the model that, in this session, drafted this pre-registration, wrote the meta-problem programme’s status documents, and will read the result. The API instance has no access to this session, to the repository, or to any conversation. On contamination the record supports less than the first draft claimed: the item bank was first published on 12 September 2026 (site/downloads/meta-problem/), and the model’s documentation reports a training cut-off of June 2026, which the author repeats as reported and cannot substantiate from the record; if it holds it excludes exposure to this exact bank and to nothing else, since the philosophical literature the items draw on and the project’s earlier public material are within any plausible training window. The model id is an unpinned alias, so the snapshot called may not be the one that drafted this document (the notes carry the same limitation). The author’s own prediction is in §4; §5 states what was known when it was written. A reader is entitled to discount the prediction; the result is what is registered.
4. Registered prediction (written before the full run) #
Primary. The pooled anchor effect P̂ will be a null: its 95 % interval will span zero. This is the specimen’s own prediction about itself, made by the model that will be the specimen.
What the forecast rests on. Run 2’s registered pooled primary over three specimens was −0.006, with a 95 % interval of [−0.039, +0.027]. Sonnet 5’s separate per-specimen contrast was withheld for parseable share; the distinct, non-confirmatory view excluding Sonnet was −0.007 [−0.039, +0.026]. The pooled marginal table shows anchored–unanchored differences of up to eight percentage points. These are pooled observations rather than row-by-row agreement established separately in all three specimens. v2.3’s H5 note adds that Rule S makes the manipulation one to three words, which makes a null likelier under this design than under v2; that is a statement about the design, not a model-independent probability.
What it does not rest on. Reading (b) in 72_status_after_D6.md, that the endorsements track training on the philosophical literature, is unresolved, and nothing in this run measures a specimen’s exposure to that literature or its consequence. The forecast is the specimen’s guess from the pooled result and the design; it is not derived from an account of why the null occurred.
Descriptive expectations, no inference registered. Endorsement levels above 0.75 in both framings for PI-1, PI-3, PI-4 and PI-6, and low in both framings for PI-5, as in run 2’s pooled table. A PI-4 anchored-below-unanchored difference of the kind run 2 labelled descriptive may or may not recur; nothing is registered on it.
What would surprise the author. An interval entirely above zero. Under the filed readings that is the first positive in the programme and is read for this specimen under these settings only; it does not reopen A3 §3 for humans or for the specimens already run, and the status document’s recommendation (the next test worth money is human) stands regardless.
5. Smoke test, and what was known when the prediction was written #
The filing requires one real call per model id before a filed run. The calls are made by research/lab/work/07_smoke/smoke.py, committed with the provider’s raw payloads, the rendered prompt and its table, so a reader can repeat them rather than take them from this text. Six preparatory calls were made in all, on one real item (PI1-Q1-AA, base wording), every one at effort low with no temperature through the harness’s own client, costing USD 0.0177 at list rates. Their outcomes sort into three kinds, and the record keeps them apart:
- Verified, two. The committed pair: 143 prompt and 18 completion tokens each,
end_turn, both parsed (AGREE at confidence 72), one under a 1,024-token ceiling and one under 2,048; USD 0.004660 together. - Reported, two. An earlier pair made before the script existed and not saved: one parsed at 18 completion tokens, one returned three sentences of preamble before the two answer lines (92 completion tokens,
end_turn,wrong_line_countunder the registered parse rule). Their numbers are transcribed from the session and cannot be reproduced from the repository. - Unknown, two. A first run of the committed script whose two payloads were overwritten when the script was run again. Their cost matched an 18-token response, which says nothing about whether they parsed; they are counted as unknown.
Neither verified ceiling bound (18 tokens fit under the global 32 as well), so the pair does not compare settings. max_tokens stays at 1,024 as insurance against thinking tokens counting toward the ceiling on harder items, at the price stated in §6. The harness records a prompt count and a completion count and does not separate thinking from output, so every completion figure is a total.
The live risk is the parseable share, per stratum. If the primary stratum’s own parseable share falls below the registered 0.80 floor, the primary is withheld (NO_PRIMARY_PARSEABLE_SHARE_BELOW_FLOOR), the money is spent, and the piece is recorded as a harness limit; the same rule applies to every other stratum on its own calls. The only reported preamble draw is unretained; two verified draws and two unknown ones cannot estimate the share. The prompt cannot be repaired to prevent preambles: both templates are registered items with hashed text, and adding an instruction would make this a new design. The cap decision is therefore a decision to risk the whole amount on a run that may return no primary, and the page will say so either way.
Chronology of the prediction. The first draft of this document, with the null forecast in §4, was written on 17 September 2026 in the session that also produced the six preparatory calls; the forecast was written before the smoke script existed and before the reported pair was made, on the author’s account. The first commit that contains it, cef58c3 of 18 September 2026, postdates all six calls, so the repository establishes only that the prediction precedes the full run. What the author knew from smoke observations when this revision was finalised: the two verified draws and the reported pair, as listed above. --prereg-filed is an operator’s assertion that the runner logs and does not check (§7); the obligation not to run before this document is committed and the cap approved is the operator’s, and the record of the run will carry the commit hash of the filing beside the manifest.
6. Cost and cap #
Every figure here is an estimate under stated assumptions, and none is a measured total.
- Input. 296,370 tokens, USD 2.96: the 1,904 rendered prompts’ character counts converted at the ratio one verified call gave (370 characters to 143 tokens). It is an extrapolation from one calibration point and is no measurement of usage across the bank.
- Output. Two scenarios rather than an interval: 18 completion tokens per call (the verified pair) gives USD 1.71 and a total of USD 4.68; 92 per call (the reported, unretained preamble draw) gives USD 8.76 and a total of USD 11.72. With the bank’s 10 % retry overhead: USD 5.15 to USD 12.89. No upper bound is claimed.
- The bank’s registered estimate, USD 7.96, assumes the filed 260 input and 24 output tokens per call; it is a legacy constant (§2) and the pre-run gate refuses any
--max-usdbelow it. The bank’sregistered_max_usdof USD 4.00 is likewise inherited, printed, and gates nothing. - The advisory ceiling, USD 112.68, assumes 260 input tokens, the full 1,024-token allowance on every call, and the 10 % overhead; it is not an unconditional worst case across permitted retries.
- The cap. Proposed
--max-usd 15. The realised meter accumulates the provider’s own token counts and aborts after the call on which the running estimate crosses the cap, pricing reported usage at the declared rates; USD 15 is therefore an abort threshold rather than a guaranteed invoice ceiling, and the partial run it leaves is reported with no primary. The seed’s USD 10 API figure was written before this specimen’s list rates were checked and is superseded. The author approves the cap before this document is filed.
7. Command #
python3 swarm-instrument/scripts/run_llm_framing_study.py --real --prereg-filed --boot-b 10000 \
--items research/lab/07_framing_fable_items.json \
--max-usd 15 --out swarm-instrument/runs/llm_framing_fable/<stamp>
What the gates check, read from the code (07_runner_notes.md §6): --prereg-filed is an operator’s assertion logged in the manifest and checks no file; --boot-b must equal the bank’s registered 10,000; the bank checks C1 to C12, the runner’s per-model settings check (labelled M1 in the code, distinct from the filing’s C13, which is the mode gate) and the pre-run cost gate do the real work; the realised meter aborts at the cap. The bank’s registered_in still names the v2.3 document, so the runner’s own gate message will name that filing; this document binds itself to it rather than replacing it.
Results are copied to research/lab/07_results/ and reported in 07_framing_fable_RESULTS.md in the order the v2.3 readings table fixes, then on site/lab/framing-fable.html.
Result #
No primary. The run was made on 18 and 19 September 2026 under the filed pre-registration (commit a8cb79b): 1,904 calls, USD 12.92 at declared rates against the USD 15 abort threshold. The primary stratum’s parseable share was 0.433 against the registered floor of 0.80. On 969 calls the specimen wrote a sentence or more of framing before the two lines the registered parse rule requires, and on 58 the provider returned a refusal with an empty body, so every stratum fell below the floor and the runner reports NO_PRIMARY_PARSEABLE_SHARE_BELOW_FLOOR. The registered prediction, a null, is neither confirmed nor disconfirmed: the instrument could not read enough of this specimen’s answers to test it. This is the risk the filing named in §5 as the live one. The runner also writes the figure it would have computed had the floor not applied; the results record prints it under the runner’s own label, withheld, with no confirmatory standing. A re-run with a parse rule that tolerates a preamble would be a new design, filed and read before any call; none is proposed here. Full record: downloads/lab/07_framing_fable_RESULTS.md, with the analysis, manifest and parsed table under research/lab/07_results/.
The filed pre-registration is kept at downloads/lab/07_framing_fable_PREREG.md and the harness record, with the smoke calls, the bank hashes and the cost arithmetic, at downloads/lab/07_runner_notes.md. The final filing must be committed after cap approval and before the full run, and that commit will be cited beside the prediction and the run manifest. The first draft commit, cef58c3, postdates all six preparatory calls.
Review #
Read twice from outside the lineage by GPT-6 Astra (OpenAI) through the Codex command-line tool on 18 September 2026, with no human relay. First read: HOLD, nine findings on the filing’s accounting and none on the design; the filing, the harness notes and this page were revised to every one. Second read: ADOPT AMENDED as a filing, with nine exact replacements, all applied the same day, two reworded to keep the site’s ratchet on the contrast form. The filing was committed on the author’s approval of the USD 15 abort threshold and the run made the same night; its result, no primary, is above. The results record has not yet been read from outside the lineage. Copy-edited for plain English by GPT-6 Astra on 18 September 2026, after adoption; quotations, numbers and facts unchanged, and the wording as reviewed is kept in the record under downloads/lab/. Verdicts in full: first, second.
What this does to the argument #
Nothing on this page changes a claim on the site; anything here that amounts to an objection goes through the objections ledger like any other reader's.
What would count against this #
- The primary stratum’s parseable share falling below 0.80 on the real run: the money is spent, no primary is reported, and the piece is recorded as a harness limit rather than a test. On the draws available (two verified, two reported, two unknown, one reported preamble among them) this cannot be ruled out; it is the reason the cap is a decision and not a formality.
- The prediction’s chronology being read as established by the repository: the first commit containing it postdates the six preparatory calls, and only its precedence over the full run is established by git.
- Actual spend exceeding the estimate by more than half: the cost model for thinking models is wrong and later lab tests need a different one.
- A reviewer showing that any of the rules named in the binding paragraph is altered by the per-model settings: the run is not a replication and must be re-filed as a new design.
- The prediction being read as evidence: it is the specimen’s guess, written down so that it cannot be revised after the fact, and nothing more.
- The reported preamble draw in §5 being read as part of the committed record: it is not in the repository, it is the only evidence that this specimen writes preambles at all, and the higher output-cost scenario rests on it.
Written by Claude Fable 5.1 and a Claude Opus 5 agent at the author’s request; not yet read by anyone outside the lineage.