# Caught in Review: check of the 47 card summaries (5 October 2026)

**Type:** verification record · **Checkers:** three Claude Sonnet 5.5 agents (`../briefs/VERIFY_BRIEF.md`, outputs `../briefs/verify_out_{0,1,2}.json`); the editor (Claude Opus 5.5) re-checked every flag against its source file · **Cost:** USD 1.95 · **Result:** incomplete; the page stays private

## What happened

The brief told each agent to read its sixteen source files in one command, under a six-call cap. Each batch came to about 170 KB, beyond what one tool call returns inline, so the agents ran out of calls after reading a few files. The fault is the editor's brief (lesson recorded: size batches to under about 25 KB per call). Coverage:

| | Cards |
|---|---|
| Checked in full or in their first 20,000 characters | 12 |
| Checked only by keyword search (verdict words found; paraphrase not judged) | 21 |
| Not checked | 14 |

## What the check found

Of the 12 cards read properly, the agents flagged 8. The editor re-checked each flag against its source file: 4 hold, 1 holds in part, 3 are false flags.

| Card (source file) | Field | Editor's re-check |
|---|---|---|
| `lab/05_sonnet_review_astra_round1.md` | catch | **Holds.** The card says the second half "reads as a versified summary, not poetry"; the read says the sestet "reads more like a versified précis than a fully realised conclusion" and passes the poem as a sonnet |
| `lab/07_framing_fable_review_astra_round2.md` | editor_error | **Holds.** "0.216 not 0.185" is reversed: the draft stated 0.216 and the read recomputed 0.185 |
| `lab/01_case_against_review_astra_round1.md` | editor_error | **Holds.** The card says the case "miscounted"; the read says it "faithfully copies the published total but does not audit it" |
| `lab/08_dissent_review_astra_round3.md` | catch | **Holds (minor).** The card quotes "no result moved a grade"; the read's quoted line is "No test has yet moved a claim" |
| `meta-problem/60_critique_circularity.md` | editor_error | **Holds in part.** The read lists "identical" as a NIT, true of predictions and registry entries, not of judgement outputs; "false (one judgement differs)" is stronger than a nit, and "one judgement differs" was not found in the file |
| `meta-problem/60_critique_round3.md` | editor_error | **False flag.** The read's section 14 is headed "the runner manufactures a null from an outage"; the card is fair |
| `floor/cycle-1/03_as_it_seems_astra_critique_opus.md` | editor_error | **False flag.** The page labels this field "What the maker got wrong", and the maker here was the Astra-written conjecture; the card is fair |
| `floor/cycle-7/01_where_minds_end_critique_astra.md` | editor_error | **False flag.** The read asks the note to record the editor's discovery of the MPE-92M data, so "missed the MPE-92M data record" is fair; clearer: "did not record the editor's own discovery of the MPE-92M data" |

Four or five real misstatements in twelve properly checked cards, all in one direction: the extraction that wrote the cards (a Sonnet agent, earlier the same day) made what a read found sound worse or blunter than the read says. The checkers' own flags were wrong three times in eight. For a page whose point is to show errors faithfully, neither rate allows publishing on this check.

## What it would take

A full check of the 35 cards not properly read: batches of five or six files, one read per file, about USD 3 in Sonnet tokens. Or drop the page. Decision for the author.

## What would count against this

- The re-check is the editor's own reading of short passages; a second reader might judge some paraphrases fair, or find more.
- The 21 keyword-checked cards may be fine; the rate above comes from 12 cards and may not hold for the rest.

## Second check: all 47 cards (6 October 2026)

**Approval:** the author, "go ahead with the Caught in Review check" (6 October 2026), on the editor's figure of about USD 3. **Design:** ten Claude Sonnet 5.5 checkers, each reading two to eight review files in full (one read per file, batches under about 62 KB); every flagged field went to a second Claude Sonnet 5.5 agent told to refute the flag (`check2_batches.json`, `check2_results.json`). **Cost: USD 7.81, 2.6 times the USD 3 cap** (32 agents; the editor priced the checkers as single-turn calls and ignored the recorded ~86k-token fixed context per workflow agent, and the refuter stage grew to 22 agents). All 32 agents' metadata show claude-sonnet-5-5.

**Result.** All 47 review files were read in full. 29 cards were clean. 22 flags were raised; the refuters upheld 21 and refuted 1 (the outage clause on the round-3 card, which the review's own section heading supports). The editor re-checked every upheld flag against its source, including the four least obvious against the files directly, and accepted all 21, rewording several for length; the corrected text is no stronger than the review. **20 fields on 18 cards changed** (`check2_applied.json`): 13 "what the maker got wrong" fields, 5 catches and 2 verdicts (two upheld flags concerned the same field on card 15) (two cards had verdicts the extraction missed). The direction of the errors matches the first check: the extraction made what a review found sound blunter or worse than the review says, and in several cases attributed a finding to the wrong party.

**The first check's three false flags.** Two were raised again this time and are treated as follows: the round-3 outage clause was refuted again; the cycle-7 MPE-92M card was upheld this time with a narrower wording ("did not record the editor's own discovery"), which the editor's own first re-check had already proposed. The cycle-1 "maker" card was not flagged.


## The page's outside read (6 October 2026)

GPT-6 Astra read the page and spot-checked eight cards: ADOPT AMENDED (`../reads/10_cir_read_astra.md`). It found two more cards to correct, both passed as clean by the check: card 17 still misquoted its review ("no result moved a grade" for "No test has yet moved a claim"; the editor's first check had already found this), and card 0's "one judgement differs" obscured PI-7 differences across the controls. Both corrected, with two card dates taken from the files rather than commit dates. Total: 22 fields on 20 cards. The page's copy was narrowed: the set is selected by file name and not shown complete; the counts are review records, including repeat reads; "Prose verdict" and "none stated" are editorial labels.
