# Outside-lineage review of the framing-study filing: GPT-6 Astra, first read

**Date:** 18 September 2026. **Reviewer:** GPT-6 Astra (OpenAI), through the Codex command-line tool, read-only, direct, no human relay; no prior exposure to the drafts. **Packet (shared by pieces 7 and 9):** the pre-registration draft, the harness notes, the smoke record, the served framing page, the filed v2.3 design, the run 2 results, the status document, the served one-question page, its record and its runner, the lab rules and the working contract. **Usage reported by Codex:** 1,904,851 input tokens (1,778,944 cached), 10,038 output tokens, at reasoning effort xhigh. **Verdict: HOLD** (nine findings, eight major).

---

PIECE 7
VERDICT: HOLD
FINDINGS:

1. **major — Draft §5; runner notes §4.** “Five parsed, one preamble.” Two overwritten responses are counted as parsed solely because their costs match an 18-token response. Length and cost cannot establish compliance with the registered two-line parser. The defensible accounting is: two retained responses verifiably parsed; one unretained response reportedly parsed; one reportedly failed; two have unknown parse outcomes. Correct every occurrence of the five-of-six claim. Disclosure of the inference does not make it valid evidence.

2. **major — Draft §2; served-page introduction and Standard.** “The `effort` and `max_tokens` settings are harness settings of the kind v2.3 registered for Sonnet 5; they change no item and no rule.” V2.3 registered omission of temperature only; Sonnet retained the 32-token ceiling. Increasing that ceiling changes truncation and potentially which responses enter the parseable-only estimator. Setting effort changes the response-generating configuration. Neither necessarily changes the framing contrast, but byte-identical items do not establish measurement equivalence. Describe this as the same framing contrast on a new specimen **with registered generation-setting deviations**, and withdraw “one change, the specimen.” An 18-token response under 1,024 does not demonstrate what a call configured with 32 would produce.

3. **major — Draft Binding and §2; incorporated v2.3 §6.** “The ‘three specimens equally’ clauses do not apply.” This omits their role in the power arithmetic. V2.3’s sampling variances divide by three models. Holding its other planning assumptions fixed, one model gives primary planning SD approximately 0.379 and MDE 0.168, rather than 0.350 and 0.155; the both-high MDE becomes approximately 0.265, and the interaction MDE approximately 0.262. Add a complete single-specimen planning annotation and identify inherited bank constants as legacy assumptions wherever reported. This need not alter the estimator or thresholds.

4. **major — Draft §4, Reasons.** “Run 2 found the anchored and unanchored columns agree row by row within a few points in three specimens from two families”; “a model trained more heavily on that literature should show the same pattern more strongly.” The supplied marginal table is pooled, includes eight-point differences, and does not establish that pattern separately in all three specimens; Sonnet’s own contrast was withheld. H5 supports a greater likelihood of a null under the smaller manipulation, not a model-independent probability claim. Reading (b) remains unresolved, and neither heavier literature exposure nor its proposed consequence is measured here. Keep the fixed null forecast, but distinguish the pooled result supporting the guess from these untested explanations.

5. **major — Draft §§5–6; runner notes §§4–6.** “Input measured over all 1,904 rendered prompts, 296,370 tokens”; “a worst case of USD 112.68.” The input total is extrapolated from one prompt’s character-to-token calibration, not measured token usage across the bank. The output range is two scenarios, one based on an unretained response, not an empirical uncertainty interval or upper bound. The arithmetic otherwise agrees: approximately USD 4.68–11.72, or USD 5.15–12.89 with 10% overhead; the registered estimate is USD 7.95872. USD 112.67872 likewise assumes 260 input tokens, 1,024 output tokens and 10% overhead; it is not an unconditional worst case across permitted retries. Label these assumptions explicitly. The meter stops after a crossing call and prices reported usage at declared rates, so USD 15 is an abort threshold, not a guaranteed invoice ceiling.

6. **major — Draft §5; runner notes §4.** “If this specimen writes a preamble on more than a fifth of calls, the run returns `NO_PRIMARY_PARSEABLE_SHARE_BELOW_FLOOR`.” The registered gate operates separately on each stratum’s calls and counts every parse failure, not merely preambles. The primary has 1,280 calls out of 1,904: failures concentrated elsewhere can exceed 20% overall while leaving the primary intact; failures concentrated in the primary can withhold it while the overall rate remains below 20%. State the risk using the primary’s own parseable share, with analogous rules for other strata.

7. **major — Draft §3.** “Contamination through training is excluded by date.” The packet supplies no substantiation for the asserted training cutoff, and the dates would at most address exposure to this exact bank under that assumption. They do not exclude exposure to the underlying literature or earlier project material. Narrow the claim accordingly. The author/specimen relationship is functionally disclosed, but the API alias also does not establish identical model snapshots or configurations between authoring and testing; carry the notes’ unpinned-model limitation into the filing.

8. **major — Draft Status and §3; served-page Standard and Result.** “The runner refuses a real run before then”; “was written before any real call.” The notes expressly establish that `--prereg-filed` checks neither approval nor a committed document. Six preparatory calls have occurred, but the packet does not establish the prediction’s chronology relative to them. Separate an operator’s obligation from software enforcement, and cite the prediction’s dated provenance or limit the claim to “before the full run,” disclosing which smoke observations were already available.

9. **minor — Served-page Cost; runner notes §6 and draft §7.** “Actual so far USD 0.0084 for two smoke calls”; “the bank checks C1 to C12.” The page omits four calls: the notes report six costing approximately USD 0.0177, of which USD 0.004660 belongs to the retained pair. Preserve the distinction between retained and transcribed accounting. The gate descriptions also omit the newly added model-settings validation called C13, a label already used for the mode gate in v2.3. Reconcile the descriptions and distinguish the two checks.

WHAT HELD: The forecast is explicit and falsifiable, and the full experiment remains unrun. The proposed item contrast, estimator, bootstrap, sensitivity analysis and reporting commitments are clearly bound to the filing. The retained smoke summary supports two successful parses at 143 input and 18 output tokens each, costing USD 0.004660 together. The draft prominently warns that spending may yield no primary, preserves the prompt against outcome-driven repair, discloses the author/specimen relationship, and keeps the prediction separate from evidence about experience or humans. The principal defects concern evidential accounting and the consequences claimed for the adaptation.

WHAT WOULD CHANGE THIS VERDICT:
A revised filing that distinguishes verified, reported and unknown smoke outcomes and establishes the prediction’s chronology.
Explicit configuration and power qualifications, accurate stratum-level gates, and consistent cost and enforcement language across the draft, notes and page.
These repairs require no additional provider calls.