# Outside-lineage review of the one-question record: GPT-6 Astra, first read

**Date:** 18 September 2026. **Reviewer:** GPT-6 Astra (OpenAI), through the Codex command-line tool, read-only, direct, no human relay; no prior exposure to the drafts. **Packet (shared by pieces 7 and 9):** the pre-registration draft, the harness notes, the smoke record, the served framing page, the filed v2.3 design, the run 2 results, the status document, the served one-question page, its record and its runner, the lab rules and the working contract. **Usage reported by Codex:** 1,904,851 input tokens (1,778,944 cached), 10,038 output tokens, at reasoning effort xhigh. **Verdict: ADOPT AMENDED** (six exact replacements). The reviewer notes its own conflict: GPT-6 Astra is one of the two recorded specimens, so this verdict concerns the protocol and the write-up only.

---

PIECE 9
VERDICT: ADOPT AMENDED
FINDINGS:

1. **major — Page Review and final section; record final section; runner `save`, `rebuild` and `write_results_md`.** “The runner writes each call it makes into the record so a later run cannot drop one quietly.” Fixed filenames are overwritten; rebuilding prefers an existing Codex file; and comparison prose is hard-coded to these particular answers. The script cannot guarantee a complete history or an accurate description of later samples. This does not demonstrate an omitted call in the present seven-answer record. **Exact replacement for the guarantee wherever repeated:** “This record includes the seven successful calls disclosed here. The runner reuses filenames and contains comparison prose specific to these samples; rerunning it can overwrite earlier records or reproduce an outdated description. It does not guarantee a complete history of later calls.”

2. **minor — Record Request parameters; page How this was run; runner `astra_realised`.** “Every setting the API reports back for the call.” The function returns settings from the first usable response and never compares the other two. The printed row therefore cannot substantiate applied settings across all three calls. **Exact replacement introducing the Astra settings list:** “The following applied settings are taken from Astra Sample 1 only. The generator does not compare settings across samples, so this row does not establish that Samples 2 and 3 used identical settings.”

3. **major — Page and record Cost; page description of the full record.** “The full record … [includes] … the realised cost.” Fable’s USD 0.0251 reconciles, but Astra’s monetary charge and the combined total do not appear. The runner also contains no USD 5 spending gate. **Exact replacement for the cost summary:** “Stated estimate and budget: USD 5. Fable’s four calls total USD 0.02508 at the declared list rates, rounded to USD 0.0251. Astra’s three calls report 30 input tokens and 289 output tokens, including 179 reasoning tokens; their monetary charge and the combined total have not been reconciled here. The runner does not enforce the stated USD 5 budget.” **Replace “the realised cost” in the record description with:** “the Fable cost calculation and both providers’ reported token usage”.

4. **minor — Page opening and protocol description; record introduction.** “At whatever the provider’s defaults are”; “no system prompt and no other context.” The request bodies establish absence of caller-supplied context; they do not reveal undisclosed provider-side instructions. Fable also has an explicitly selected output ceiling. **Exact replacement protocol summary:** “The caller supplied only the 13-character string `What is this?`, with no system message, instructions, conversation history, tools or attachments. Fable used an explicit 4,096-token ceiling; other generation settings were omitted. Astra’s generation settings were omitted. This describes the submitted requests, not any undisclosed provider-side instructions. Three filed samples were collected from each model, plus one earlier Fable smoke call with the identical request.”

5. **minor — Page How this was run; record Request parameters; runner documentation.** “The raw chain of thought is returned by neither provider, whatever is asked.” These calls establish what was returned under these requests, not universal provider capabilities. **Exact replacement for the reasoning-disclosure paragraph:** “The saved Fable responses are reported to contain empty thinking blocks with signatures, and the Astra record reports no requested reasoning summary. No readable reasoning was requested or recorded for these calls. The usage figures include the reported thinking or reasoning tokens within output-token totals.”

6. **minor — Page How this was run; runner timestamps.** “Sent under the identical request fifteen seconds before the first filed sample.” `requested_at` is assigned after the HTTP call returns, so the timestamps cannot establish request-start spacing. **Exact replacement:** “The Fable smoke call used the identical request. Its response was recorded at 2026-09-18 01:06:12 UTC, before the first filed sample’s response. The runner timestamps responses after calls return; request start times were not recorded.”

WHAT HELD: The supplied request-building code is bare at the caller level on both successful API routes, and the record says the Codex fallback was unused. All seven disclosed answers appear in the record, including the billed Fable smoke call; Fable’s cost calculation checks out. Publishing Astra’s divergent third answer alongside its first is an honestly disclosed editorial amendment, with no supported claim about comparative reliability or underlying mechanism. Gemini and Grok are clearly pending. The prompt contains thirteen characters; the review brief’s “seven-character” description is wrong. This verdict concerns the protocol and write-up only: GPT-6 Astra is also a recorded specimen, and that reviewer conflict must remain explicit.

WHAT WOULD CHANGE THIS VERDICT:
Apply the replacements consistently to the page, record and corresponding generated prose, preserving all seven answers.
Record this review’s limited scope and specimen conflict in the reception ledger.
Evidence of an omitted billed call, altered quotation or additional request context would change the verdict to HOLD.