# TypeSafe on the project's questions: what Jev returned, against what the editor expected

**Type:** record · **Date:** 21 September 2026 · **Status:** run once, as pre-registered in `00_DESIGN.md`; nothing repaired; the pre-registered rules for changing a next step are applied below · **Authors and models:** editor Claude Fable 5.1; model queried TypeSafe Jev 1.13.0 (`jev-latest`), twelve requests · **Standard:** the design's · **Cost:** 71,546 input tokens at USD 0.042 per million, about USD 0.003; the author's key, added to a gitignored file and never printed

## Set A, next steps: Jev puts the dated-decision rule first

| option | Jev | editor expected |
|---|---|---|
| dated-decision rule | **0.67** | low |
| go live with the review | 0.19 | 0.15 |
| correct P1's falsifier | 0.10 | 0.35 |
| S1 reconstruction | 0.01 | 0.25 |
| panel desk tier, derivation audit, fourth-family judge, none | 0.01 each or less | spread |

Confidence 0.62. The pre-registered rule (an item other than the P1 correction first, confidence above 0.5) fires: the editor re-orders the list with the dated-decision rule first and says so. Jev's reasons are not available; its choice is consistent with the state's own emphasis on a machinery that has never lowered a grade, which the editor wrote into the state and had ranked fourth.

## Set B, the three findings: all three hold on Jev's reading, two more strongly than the editor expected

| question | Jev | expected |
|---|---|---|
| B1 causal set theory satisfies P1's falsifier as worded | 0.78 | 0.75 |
| B2 S1's strong reading, as Part I argues it, depends on P2 | 0.91 | 0.65 |
| B3 the unchanged grades indicate the machinery cannot lower one | 0.55 | 0.45 |

No number below 0.30, so no third read is triggered before a correction is filed. B3 near 0.5 is what Astra's check said: the pattern warrants an explanation, not a verdict of incapacity.

## Set C, a fourth-family blind judge: Jev never chose Astra

Per triple, Jev's choice of best output with its lineage resolved, its probability for each arm, and the sum of its five yes/no judgements per output (out of 5):

| triple | Jev's best | conf | P(Fable) | P(Opus) | P(Astra) | yes/no sums F / O / A |
|---|---|---|---|---|---|---|
| QG | Fable | 0.49 | 0.66 | 0.22 | 0.12 | 4.08 / 4.22 / 3.80 |
| CAT | Fable | 0.14 | 0.43 | 0.32 | 0.25 | 4.25 / 4.11 / 4.02 |
| PHY | Opus | 0.45 | 0.25 | 0.63 | 0.12 | 4.29 / 4.36 / 4.03 |
| PHI | Opus | 0.30 | 0.41 | 0.53 | 0.06 | 4.26 / 4.32 / 3.98 |
| HIS | Opus | 0.21 | 0.19 | 0.47 | 0.34 | 3.87 / 4.00 / 4.06 |
| PRE | Opus | 0.20 | 0.32 | 0.47 | 0.21 | 4.02 / 4.13 / 4.09 |
| SYN | Fable | 0.65 | 0.77 | 0.12 | 0.11 | 3.93 / 4.01 / 3.78 |
| AUD | Fable | 0.48 | 0.65 | 0.28 | 0.07 | 4.19 / 4.11 / 3.72 |

Best counts: Fable 4, Opus 4, Astra 0. Mean probability of being best: Fable 0.46, Opus 0.38, Astra 0.16. Pairwise, ranking by Jev's probabilities: Fable against Opus 4 to 4; Fable against Astra 7 to 1, p = 0.07; Opus against Astra 8 to 0, p = 0.008.

The pre-registered rule (Set C separates the arms 8 to 0 in either family's favour) fires for Opus against Astra. So the record now says: the partisanship reading in `20_comparison.md` is weakened. A judge from a third provider family, blind, with no argument and no stake, ranked the two Claude arms the way the two Claude judges did and never ranked Astra first, where the Astra judge had ranked Astra first eight times. That does not prove the Claude judges were right about quality; it does mean their ranking is no longer explained by family alone, and the Astra judge's is now the one no other judge shares. Astra's own check had already conceded that its judge "applied an uneven standard".

Two caveats the design named. Jev's yes/no judgements barely discriminate: every output scores between 3.7 and 4.4 of 5, mostly at 0.8 to 0.9 per question, so the signal is in the Choice, not the Nouls, and a Choice can track surface features (length, form, confidence of tone) as well as substance; the Claude outputs were the longest. And Jev is a fast-judgement model outside its training domain here; its confidences on the Choice run from 0.14 to 0.65, low where the outputs were close.

## Set D, outreach and risk

| question | Jev | expected |
|---|---|---|
| D1 first person: Kammerer / Chalmers / Frankish / Seth / Baron | **0.62** / 0.34 / 0.02 / 0.01 / 0.01 | 0.40 / spread / 0.30 |
| D2 form: disclosed AI-sent / author writes / public post / none yet | **0.86** / 0.14 / 0.00 / 0.00 | 0.25 / 0.50 |
| D3 a disclosed AI-sent message gets a considered reply | 0.42 | 0.30 |
| D4 objection 17 read as honesty machinery, not collapse | 0.39 | 0.55 |
| D5 the review read as honesty machinery, not noise | 0.48 | 0.60 |

The rule on D2 (author-writes clearly ahead) does not fire; Jev leans the other way. The rules on D4 and D5 (below 0.35) do not fire, but D4 at 0.39 is close: on Jev's reading a hostile reader is nearly as likely to read objection 17 as the project collapsing as to read it as the machinery working. The editor takes that as a caution on how item 2 is written, not on whether: the objection must carry the confidence policy (the 0.8 was never the strong reading's) in its first lines.

## Set E, the wager

Diagnosis correct 0.47 (editor 0.50); spacetime and consciousness two appearances of one self-referential structure 0.42 (editor 0.20; the site 0.3); current AI language models conscious 0.09 (editor 0.05). Nothing moves; registered as such.

## What changes in the proposed next steps

1. The dated-decision rule moves to first, by the rule in the design. It costs nothing and it is what a third family's judge, a fourth family's judge and the outside check all pointed at.
2. The partisanship reading is weakened, as the design said it would be on an 8 to 0. The lab still needs a reasoning judge from outside both families for anything that matters; Jev is a cheap tiebreak, not that.
3. Objection 17's framing is flagged (D4). Otherwise the list stands.

## What would count against this

- Jev's Choice tracking length: the Claude outputs were longest and Jev never chose Astra. A rerun with the outputs truncated to a common length, or with the shortest output in each triple padded, would test it; not run.
- The states being the editor's summaries (sets A, B, D): committed under `state/` so a reader can see what Jev was told.
- The expectations being anchored on the editor's conclusions: they were, and Jev departed from them in three places (A, D2, E2).
