# Outside-lineage review of the framing-study filing: GPT-6 Astra, second read

**Date:** 18 September 2026. **Reviewer:** GPT-6 Astra (OpenAI), through the Codex command-line tool, read-only, direct, no human relay. **Packet:** the revised draft, the revised notes, the smoke record, the served page, the first-read verdict, the filed v2.3 design, the run 2 results, the status document, the relevant runner functions, the lab rules and the working contract. **Usage reported by Codex:** 2,304,053 input tokens (2,171,904 cached), 13,038 output tokens, at reasoning effort xhigh. **Verdict: ADOPT AMENDED** as a filing (two first-read findings resolved, seven partly; nine exact replacements, all applied the same day across the draft, the notes, the smoke metadata and the page, two reworded to keep the site's ratchet on the contrast form). The reviewer's condition: the filing can be committed before the full run once the author approves the USD 15 abort threshold.

---

VERDICT: ADOPT AMENDED

FIRST-READ FINDINGS:

1. partly resolved — The verified/reported/unknown classification correctly preserves all six calls, but notes §4 still says “Four real draws in all.”
2. partly resolved — The draft and page register the configuration deviations correctly; notes §4 still claims equal parseable shares at 1,024 and the untested 32-token setting.
3. partly resolved — Most single-specimen arithmetic is correct, but the paraphrase-only MDE is wrong and several corresponding contrasts remain unspecified.
4. partly resolved — The unsupported literature-exposure explanation is withdrawn, but the revision incorrectly describes Run 2’s confirmatory pooled primary as a two-specimen result.
5. partly resolved — The draft qualifies the cost assumptions and abort threshold correctly; the notes retain measured-usage, bounded-range and worst-case language.
6. resolved — The operative descriptions now apply the 0.80 floor per stratum, count every parse failure and identify the primary’s 1,280 calls.
7. resolved — The cutoff is explicitly unsubstantiated, its conditional implication is narrowed to the exact bank, and the unpinned-alias limitation is disclosed.
8. partly resolved — The chronology now distinguishes author testimony from commit evidence, but the draft’s Status and the page’s Result retain the original enforcement and chronology overclaims.
9. partly resolved — The preparatory expenditure is reconciled and M1 is distinguished from C13; notes §6 still omits M1 from its consolidated bank-check description.

NEW FINDINGS:

1. **major — Draft §2 and served-page §2: incorrect and incomplete power annotation.** “The paraphrase-only stratum at 4 samples has MDE 0.216.” This remains the quad average of two differences: its sampling variance is ¼ × 4 × (0.25/4) = 0.0625, planning SD 0.418001, and MDE 0.185163. Replace that clause with:  
   “The paraphrase-only stratum at 4 samples has sampling variance 0.0625, planning sd 0.418 and MDE 0.185 against v2.3’s 0.162.”  
   Add immediately before the sentence about inherited constants:  
   “The appearance-vocabulary main effect shares the primary’s planning sd 0.379 and MDE 0.168; the v2 diagonal shares the single-level cell’s planning sd 0.418 and MDE 0.185; the base-only stratum shares the paraphrase-only planning sd 0.418 and MDE 0.185.”

2. **major — Draft §4 and served-page §4: pooled primary confused with dropped-specimen disclosure.** “The primary’s interval spans zero over the two usable specimens”; “a pooled null from two specimens.” The supplied results record identifies a three-specimen pooled primary and a separate, non-confirmatory analysis dropping Sonnet. Withholding Sonnet’s individual contrast does not itself remove Sonnet from the pooled estimator. Replace the first two sentences of “What the forecast rests on” with:  
   “Run 2’s registered pooled primary over three specimens was −0.006, with a 95 % interval of [−0.039, +0.027]. Sonnet 5’s separate per-specimen contrast was withheld for parseable share; the distinct, non-confirmatory view excluding Sonnet was −0.007 [−0.039, +0.026]. The pooled marginal table shows anchored–unanchored differences of up to eight percentage points. These are pooled observations, not row-by-row agreement established separately in all three specimens.”

3. **major — Runner notes §4, consequence 3: unsupported counterfactual and expenditure bound survive.** “1024 and the global 32 give the same parseable share”; “What it costs is the worst case: USD 112.68.” No 32-token call was made, and the advisory calculation is not a worst case. Replace consequence 3 in full with:  
   “3. **`max_tokens` 1024 is a registered precaution, not a demonstrated requirement.** The retained responses contain 18 completion tokens each, but no call was made at a 32-token ceiling, so neither its response distribution nor its parseable share is known. The 1,024-token setting provides room for completion usage, including thinking, on other items; its benefit is unmeasured. The USD 112.68 advisory scenario assumes 260 input tokens and 1,024 completion tokens on each of 1,904 calls, plus 10 % retry overhead. It is not an unconditional worst case. The realised meter aborts after a call crosses the approved threshold and does not guarantee an invoice ceiling.”

4. **major — Runner notes §5: estimated token usage still described as measured.** “This is measured from the plan the runner builds, so the spread is the real one.” The measured quantities are prompt character counts; token usage is extrapolated. Replace the input paragraph with:  
   “Applying the single calibration ratio of 143 input tokens per 370 prompt characters to the 1,904 rendered prompts gives estimated input usage of approximately 134–184 tokens per call, mean 155.7, and 296,370 tokens in total, costing USD 2.96 at the declared rate. These are character-calibrated estimates, not measured token usage across the bank.”  
   In the draft/page §2 cost-model row and notes §6, replace categorical claims that the input is “over-counted” with:  
   “The 260-token input assumption exceeds the character-calibrated estimate; actual usage across the bank has not been measured.”

5. **major — Runner notes §§5–6 and served-page Cost: scenario and cap language remains inconsistent.** “Between about USD 4.70 and USD 11.70”; “a cap of USD 15 covers it with room”; the advisory ceiling “is the other end.” These recreate the bounds withdrawn by draft §6. The arithmetic itself agrees.

   Replace “Output depends on how often the model writes a preamble” with:  
   “Output cost depends on total completion usage, including thinking and visible text; the following rows are scenarios, not bounds.”

   Replace the paragraph beginning “So: between” with:  
   “The 18-token and 92-token output scenarios give totals of USD 4.68 and USD 11.72, respectively, or USD 5.15 and USD 12.89 with the assumed 10 % retry overhead. The 92-token scenario rests on an unretained, reported response and is not an upper bound. The proposed USD 15 is an abort threshold, not a guaranteed completion budget or invoice ceiling. The seed’s USD 10 API allocation is proposed to be superseded by this approval.”

   Replace notes §6’s pre-run cost-gate paragraph with:  
   “The pre-run cost gate compares `--max-usd` with USD 7.95872, displayed as USD 7.96, calculated from 260 input and 24 completion tokens per call at the bank’s declared rates with 10 % retry overhead. It refuses a lower cap. The advisory USD 112.67872 scenario, displayed as USD 112.68, instead assumes 1,024 completion tokens per call with the same input assumption and overhead; it gates nothing and is not an unconditional worst case.”

   Replace notes §6’s “realised spend at the proposed cap is about three times that figure” with:  
   “The proposed USD 15 abort threshold is 3.75 times the legacy USD 4.00 figure; actual expenditure is not yet known.”

   Replace the served-page Cost text with:  
   “Output-cost scenarios: USD 4.68 and USD 11.72 under the stated assumptions, or USD 5.15 and USD 12.89 with 10 % retry overhead; neither is an upper bound. Proposed abort threshold USD 15, awaiting approval. Preparatory usage priced at declared rates totals approximately USD 0.0177 for six calls, including USD 0.004660 for the retained pair; the remaining accounting is reported. Build cost is recorded in the lab’s decision log.”

   Replace remaining “upper end of the cost range” references in the current draft, notes and page with “higher output-cost scenario.” Replace notes §6’s meter sentence with:  
   “The realised meter prices provider-reported usage at the declared rates and aborts after the call on which the accumulated estimate crosses `--max-usd`, retaining a partial run with no primary.”

6. **major — Draft Status and served-page Result: contradictory enforcement and chronology claims remain.** “The runner refuses a real run before then”; “The draft is committed before any real call.” The first contradicts the supplied `mode_gate`; the second contradicts the six preparatory calls and the disclosed commit chronology.

   Replace the draft Status’s final two sentences with:  
   “No full run has been made. The operator must obtain approval of the cap and commit the final filing as `07_framing_fable_PREREG.md` before the full run. The runner requires `--prereg-filed` but does not verify approval, file existence or commitment.”

   Replace the served-page Result’s final sentence with:  
   “The final filing must be committed after cap approval and before the full run, and that commit will be cited beside the prediction and the run manifest. The first draft commit, cef58c3, postdates all six preparatory calls.”

7. **minor — Runner notes §4 and smoke-record metadata: residual count and configuration labels.** “Four real draws in all: three parsed, one did not”; smoke.json describes both calls as using “the bank’s own settings,” although call 2 uses 2,048 rather than the bank’s 1,024.

   Replace the four-draw sentence with:  
   “Across six preparatory draws, two retained responses verifiably parsed, one unretained response reportedly parsed, one reportedly failed parsing, and two overwritten responses have unknown parse outcomes.”

   Replace smoke.json’s `what` value with:  
   `"lab piece 7 smoke: two retained real calls on one item, both at effort low without temperature; max_tokens 1024 and 2048 respectively"`

8. **minor — Runner notes §6: incomplete check inventory.** “The bank checks (C1 to C12, `validate_bank` and `substitution_check`) do the real work.” M1 is explained elsewhere but omitted here. Replace that sentence’s opening with:  
   “The bank checks C1–C12 (`validate_bank` and `substitution_check`) and the additional per-model settings check M1 run before any provider call. M1 is distinct from v2.3’s C13 mode gate.”  
   Retain the subsequent bank counts and exit consequence. The supplied runner docstring supports the naming distinction; the excerpt ends before the actual M1 block, so this packet does not independently demonstrate its implementation.

9. **minor — Runner notes §7, notes §6 table reference and served-page footer: stale status labels.** “No site page”; “What changes, and only this”; “not yet read by anyone outside the lineage.” These contradict the supplied page, revised heading and first verdict.

   Replace notes §7 with:  
   “No full run has been made. A draft site page exists. The pre-registration remains unfiled and the cap unapproved. The corrected filing and page must agree before filing and the full run.”

   Replace the obsolete table title with:  
   “What changes: the specimen, and three generation settings named as deviations.”

   Replace the footer with:  
   “Written by Claude Fable 5.1 and a Claude Opus 5 agent at the author’s request; outside-lineage reviews are recorded above.”

WHAT HELD: The registered contrast, parsing rule, estimator, bootstrap and outcome-reporting commitments remain intact. The 1,904-call plan and 1,280-call primary agree; the retained smoke summary supports two successful parses at 143 input and 18 completion tokens each, costing USD 0.004660 together. The cost calculations and the other stated single-specimen MDEs recompute correctly. The narrowed contamination disclosure and explicit limits on prediction provenance are appropriate. These are document and metadata repairs; no additional provider calls or analytical redesign are required. This review does not independently revalidate the underlying bank hashes, raw payloads or reported offline test results.

WHAT WOULD CHANGE THIS VERDICT:
Apply the replacements consistently across the draft, notes, smoke metadata and served page.
After those amendments and approval of the USD 15 abort threshold, the filing can be committed before the full run.
Cap approval alone is insufficient to file the packet unchanged.