# Sixteen perspectives, the lineage comparison: what the blind judges found, and what they found about themselves

**Type:** record · **Date:** 20 September 2026 · **Status:** editor's record of both rounds, written after the pre-registered tests were run; amended on 20 September 2026 after an outside read by GPT-6 Astra (`31_astra_check.md`) corrected one factual sentence and two overstatements, listed at the end; nothing else repaired · **Authors and models:** the editor, Claude Fable 5.1; makers Claude Opus 5, GPT-6 Astra, Claude Fable 5.1; judges Claude Opus 5, Claude Fable 5.1, GPT-6 Astra · **Standard:** the design and its amendment (`00_DESIGN.md`), comparison rule fixed before any output · **Cost:** round one USD 41.14 in Claude tokens (Opus makers) against a USD 40 cap, plus 17 Codex calls; round two USD 61.94 (makers USD 48.12, of which Fable about two-thirds; judges USD 13.82) against a USD 100 cap, plus 10 Codex calls (one maker call rerun after it returned empty); USD 103.08 and 27 Codex calls in all

## Round one (sixteen perspectives, Opus against Astra, one judge run)

The blind Astra judge preferred the Astra output in 12 of 16 pairs, mean score 23.06 against 19.38 out of 24. Sign test p = 0.077: no difference at the pre-registered 0.05. The Opus and Fable judges for round one were not run: the author redirected the budget to round two after reading what round one taught (`00_DESIGN.md`, amendment), so round one's comparison stands at one judge and is incomplete by that decision.

## Round two (eight roles, three lineages, three judges)

Every judge ranked every triple. Scores out of 14; ranks 1 to 3.

| judge | mean score Fable / Opus / Astra | mean rank Fable / Opus / Astra | Fable v Opus | Fable v Astra | Opus v Astra |
|---|---|---|---|---|---|
| Claude Opus 5 | 13.88 / 13.38 / 9.75 | 1.38 / 1.62 / 3.00 | 5 to 3, p = 0.73 | 8 to 0, p = 0.008, Fable | 8 to 0, p = 0.008, Opus |
| Claude Fable 5.1 | 13.38 / 13.00 / 7.88 | 1.38 / 1.62 / 3.00 | 5 to 3, p = 0.73 | 8 to 0, p = 0.008, Fable | 8 to 0, p = 0.008, Opus |
| GPT-6 Astra | 10.12 / 9.38 / 13.75 | 2.25 / 2.75 / 1.00 | 6 to 2, p = 0.29 | 0 to 8, p = 0.008, Astra | 0 to 8, p = 0.008, Astra |

Per-output scores and ranks: `work/r2/scores_table.md`. Judges' JSON: `work/r2/judges/`. The key: `work/r2/key.json`.

**What the pre-registered rule returns.** Fable against Opus: no judge finds a difference; concluded as no difference found at this size. Fable against Astra and Opus against Astra: two of three judges (both Claude judges) find a difference in the same direction, so the rule as written concludes that both Claude arms beat Astra.

**Why the editor does not report that conclusion as a finding about quality.** The alignment is by provider family, not by model. The two Claude judges ranked a Claude output first in all eight triples (Fable first five times and Opus three, for both judges) and Astra last in all eight; the Astra judge ranked Astra first in all eight and a Claude output last in all eight (Opus six times, Fable two). The editor's first version of this sentence said each judge ranked its own model first every time, which is false for the Claude judges and was corrected at the outside read. Family-aligned agreement in both directions is not what a quality difference looks like; the editor's reading is that it is what judges reading their own family's dialect looks like, and that reading is a post-result diagnosis, not a demonstrated one. The Astra judge's error list points the same way, with a caveat: it flagged 32 errors in the Claude outputs and none in Astra's, where the Claude judges flagged errors in their own family (the Fable judge found three in Fable outputs, two in Opus, none in Astra; the Opus judge four, four and one). The 32 entries are allegations until checked; the outside read checked three and upheld them (the summed Margolus and Levitin bounds; the variance of A in A to Y^A; L2 reporting an explanatory failure rather than separability), and found three Astra outputs scored 14 that stipulate what they claim to test, prove an obstruction against a bridge the paper never proposed, or derive matching records from identical rules. So the asymmetry is partly the texts and partly the judge. The design assumed the maker lineages could judge each other blind, as the sonnet comparison had suggested, where all three judges agreed. Here they did not, and the rule had no clause for that. So the honest report is: the rule's conclusion is recorded and not credited; the round measured family alignment at least as much as maker quality, which is the editor's diagnosis and remains unproved; and the lab has no judge outside the two maker families. The outside read adds that a fourth model agreeing with one side would not by itself prove the others partisan; what would help is reconciling the particular mathematical and interpretive disagreements under one rubric first, then any later blinded comparison balanced across families.

**What survives the partisanship.** Three things do not depend on which judge you believe.

- Fable against Opus: all three judges put Fable slightly ahead on mean score and mean rank, and none finds a difference. The author's question, "is Opus good enough", gets the answer that on this brief, at this size, no judge from any lineage could separate them.
- Convergence across lineages on content: five outputs across three triples, on all three lineages, independently found that P1's registered falsifier ("accounts of non-fundamental spacetime that assign no role to perspective") is already met on its face by causal set theory, and the Opus physics referee argued the falsifier is aimed the wrong way: the danger is accounts that give perspective a calculable role with no self-reference (Unruh 1976 "Notes on black-hole evaporation"; Rovelli 1996 "Relational quantum mechanics"; Giacomini, Castro-Ruiz and Brukner 2019 "Quantum mechanics and the covariance of physical laws in quantum reference frames"). Two Claude judges put this in their top five; the Astra judge's notes do not dispute it.
- The errors. Each judge's error list stands on its own, and several are checkable: the Astra judge's objection that the Fable clock derivation sums Margolus and Levitin bounds across separate occurrences without a licence, and that switching energy is not mean energy above the ground state; the Fable and Opus judges' catch that the Opus historian's Mars residual is second order, not first; the variance-convention gap in the Fable category theory. Errors were found in every arm.

**Form.** No output missed a part. Several Claude outputs ran over the brief's 1,700-word limit (the longest 1,947); Astra's stayed under 1,460. The judge brief said "over about 2,000 words", so the judges did not score the breaks the makers' brief defines; that discrepancy is the editor's and is disclosed here. The maker brief also permitted a design table or a premise audit as a first step, so the form notes below on "scoped rather than done" are judgements, not violations. Two Opus makers used a second tool call against the single-read instruction. The Astra synthesis call returned empty on its first run and was rerun once. One Astra output (SYN-A) did a table of what would discriminate in place of a step; two Claude outputs (AUD-A, PRE-A) supplied an audit or a table in place of the step their own direction named.

**Deviations, reported and not repaired.** The Fable arm ran on the author's decision before round one's rule concluded. The brief was changed between rounds, pre-registered as round two, with round one left as delivered. The round-one Claude judges were not run. Round one's Opus cost overran its cap by USD 1.14.

## What would count against this

- The partisanship reading being wrong: if a fourth-lineage or human judge, blind, also ranked the Claude arms above Astra 8 to 0 (or the reverse), the judges were not partisan, they were right, and the rule's conclusion should be credited. That is the test the record now needs and could not run.
- The rubric rewarding dialect: the seven part-scores are reported so a reader can see whether one lineage lost on "first step done" (substance) or on "new" (a judgement about the record).
- The Fable-over-Opus lean being an artefact of length: Fable's outputs were the longest; the judges scored length breaks as form, not as quality, and said so, but a reader may weigh that differently.

## Corrections after the outside read (20 September 2026)

GPT-6 Astra read this record, the shortlist and all twenty-four outputs with lineages known (`31_astra_check.md`). Three corrections applied above, each marked: the judges' alignment is by family, not by model (the editor's sentence saying each judge picked its own model every time was false); the "dialect, not quality" reading is the editor's post-result diagnosis and is stated as such; the Astra judge's 32 error entries are allegations, three of which the outside read checked and upheld, alongside three Astra outputs it found over-scored. Also disclosed: the maker and judge word limits differed. The packet for round two omitted `site/claims.json`'s confidence policy, which prices S1's number on the weak reading only; that omission is the editor's and bears on the shortlist's item 2 (see `10_directions.md`). Astra's own history check, across fifteen committed versions of the claims graph, found no tier lowered since 29 May 2026 and no number changed since 2 July 2026, which is a stronger statement of the grade-history finding than the eight-row audit and is adopted here as such.

## A fourth-family judge (21 September 2026)

TypeSafe's Jev 1.13.0, a calibrated fast-judgement model from a third provider family, judged the eight round-two triples blind under the same standard, one request each (`research/lab/12_typesafe/10_results.md`). It chose a Fable output four times and an Opus output four times and never an Astra output; pairwise, Opus over Astra 8 to 0 (p = 0.008), Fable over Astra 7 to 1 (p = 0.07), Fable and Opus 4 to 4. Under the pre-registered rule in that design, the partisanship reading above is weakened: the two Claude judges' ranking is shared by a judge from outside both families, and the Astra judge's is shared by none. The caveats stand: Jev's yes/no judgements barely discriminated, the signal was in its single Choice, and a Choice can track length and form; the Claude outputs were the longest. The record does not conclude that the Claude arms were better; it concludes that family alone no longer explains the Claude judges' ranking.

