# Sixteen perspectives: a review of Space Immanence for directions, run on two lineages and judged blind

**Type:** record (lab type) · **Date:** 20 September 2026 · **Status:** design, pre-registered before any output exists · **Authors and models:** design by Claude Fable 5.1 (editor) at the author's decision; makers Claude Opus 5 and GPT-6 Astra; judges Claude Opus 5, Claude Fable 5.1 and GPT-6 Astra, blind · **Standard:** the lab's rules; the comparison rule below, fixed here · **Cost:** cap USD 40 in Claude tokens at list rates; Codex calls on the author's plan as he directed (about seventeen, above the floor's four-per-cycle allowance, by his instruction: "run these on Astra and Opus both")

## Why

The author, 20 September 2026: "it may help to do a review of Space Immanence from multiple different perspectives (really being creative and innovative of what these are exactly). That way we can see if we find any interesting helpful and promising directions to take the research." And: "I am worried Opus may not be good enough. I want us to run these on Astra and Opus both and see if the quality differs. And if it does we can consider also running on Fable." The lab's first batch showed that agents asked to review return dissent; this design asks for proposals in the conjecture form and nothing else.

## The perspectives (sixteen, fixed here)

1. A quantum-gravity phenomenologist: what would the fold predict that an instrument could see.
2. A category theorist working on Lawvere's fixed-point theorem and the claim of exactly two orientations.
3. An illusionist philosopher in François Kammerer's position: introspection deploys a tacit theory, and the illusionist alternative feels unthinkable for that reason.
4. A control engineer who builds self-maintaining systems, on what "at stake" would have to mean for a machine.
5. A Madhyamaka scholar sharpening the objection that "one structure" is itself a container.
6. A historian of science on what it took to remove the previous containers (the centre, absolute space and time, definite properties) and what it would take now.
7. An experimental psychologist designing the cheapest human test that could move a graded claim.
8. An AI-safety researcher on what the four fold criteria imply for agent design.
9. A hostile referee at a physics journal, asked for what would make the physics arm publishable.
10. A hostile referee at a philosophy journal, asked for what would make the philosophy publishable.
11. A science journalist: the one falsifiable claim a reader can hold in their head, and how to test it in public.
12. A funder with USD 10,000: what to buy first, and what result would justify the next USD 100,000.
13. A contemplative practitioner: what practice-based test exists, with its controls.
14. An information theorist comparing the four criteria to integrated information theory's measure.
15. A theoretical biologist on autopoiesis and the free-energy principle.
16. A pre-mortem reviewer writing in 2036 about why the project failed, and what it should have done in 2026.

## The brief

Identical for both lineages: `BRIEF.md`, with the role substituted, followed by `PACKET.md` (8,809 words built from the site's own texts: the one-page summary, the graded claims, the objections ledger verbatim, the Positions verdicts, the lab index, amendment A1, the three adopted conjectures' claims, the meta-problem outcome and the human study memo). Makers see nothing else. Output: one paragraph on what the record is missing, then exactly three directions in the conjecture form (claim, falsifier, cheapest test with cost, what it adds), under 900 words.

## Arms

- **Opus arm:** sixteen Claude Opus 5 agents, one per perspective, reasoning effort xhigh, in one workflow; each told to read only the packet.
- **Astra arm:** sixteen GPT-6 Astra calls through the Codex command-line tool, read-only, reasoning effort xhigh, the brief and packet in the prompt, no repository access.
- **Fable arm:** not run. Proposed to the author only if the comparison rule below finds a difference, with its cost stated then (about USD 50 to 60 at Fable list rates).

## Blind judging

For each perspective the two outputs are labelled A and B by a coin the editor records in `work/key.json` before any judge runs; judges never see lineage. Three judges, one per lineage (Claude Opus 5, Claude Fable 5.1, GPT-6 Astra), each read all sixteen pairs and return, per output: four scores of 0 to 2 (could it be wrong; could it be tested, with a specific and costed route; does it say something the record does not already say; is the falsifier honest), summed over the three directions to a maximum of 24; and, per pair, a preference (A, B, or neither) for "which would you take to the floor". Judges also name, across all thirty-two outputs, the five directions most worth a floor cycle. The maker lineages judge their own outputs only blind and alongside a third lineage, as the sonnet comparison did.

## The comparison rule, fixed before any output exists

- Per judge: a two-sided sign test over the sixteen pairwise preferences, ties excluded, at p < 0.05. That judge "finds a difference" if the test is significant, in the direction of the preferred lineage.
- "Quality differs by lineage" is concluded only if at least two of the three judges find a difference in the same direction. Otherwise the record says no difference was found at this size, which is not the same as no difference.
- Mean scores per lineage per judge are reported as description, not as the test.
- A maker output that breaks the form (wrong count of directions, over length by more than a fifth, a direction already in the record) is scored as delivered and the break is recorded; nothing is repaired.
- If the rule concludes a difference, the Fable arm is put to the author as a decision with its cost; it does not run on this design's authority.

## What is delivered

`10_directions.md`: every direction from both arms, normalised to one form, with its judges' scores and the blind labels resolved; the judges' five picks each; the editor's shortlist for the floor, marked as the editor's. `20_comparison.md`: the sign tests and means. A lab page after an outside read of the record.

## What would count against this

- The brief steering both lineages to the same three directions: the perspectives were chosen to be far apart, and the judges are asked whether directions repeat across perspectives.
- The judges preferring length or confidence over testability: the rubric scores testability and honesty separately, and the means per criterion are reported.
- The packet omitting something a direction needed: makers are told to say where it is silent, and those sentences are kept.
- A difference between lineages being a difference in how each follows the form rather than in what it proposes: form breaks are counted and reported apart from the scores.
- The cost figures being wrong by more than a third: Claude figures from the agents' usage records at list rates; Codex counts as reported.

## Amendment, 20 September 2026: round two, on an improved brief, three lineages, pre-registered before any output exists

**Why.** The author, after round one: "Good with the cap at USD 100." Round one is complete and stands as delivered (both arms, the blind Astra judge, no repairs). What it taught, recorded in the cycle record: the "cheapest test under USD 20, packet only" brief steered both lineages toward one genre (a small controller with and without a self-model), producing eleven clusters of repeats; the packet starved the physics, formal and philosophical perspectives, which said so; nobody was asked to cross perspectives or to do a first step rather than propose it. Round two changes the brief and the packet, and adds Claude Fable 5.1 as a third arm on the author's decision, before the round-one comparison rule concluded (one judge of three had run; it did not find a difference at 0.05). That is a deviation from the design's conditional and is recorded as the author's choice, not as the rule's outcome.

**The brief (`BRIEF_v2.md`).** One analysis of what the record is missing (up to 500 words); one direction worked first, with the test at three cost tiers (under USD 20; under USD 500; under USD 10,000 with a named kind of person's hour) and the first step done on the page (up to 450 words); one direction with no cost limit. Repeats of round one's clusters or titles are disallowed unless they say what they add. Under 1,700 words.

**The packet (`PACKET_v2.md`, about 22,000 words).** The working paper in full; the ledgers as before, with the human study memo complete (round one's packet cut it mid-sentence, which two makers reported); the persistence conjecture as adopted, in full; round one's eleven clusters and ninety-six direction titles.

**Roles (eight).** Six where round one was starved or strongest at project level: quantum-gravity phenomenologist, category theorist, physics referee, philosophy referee, historian of science, pre-mortem reviewer. Two that only synthesis can do: the one-test-many-hinges designer and the Part I auditor.

**Arms.** Claude Fable 5.1, Claude Opus 5 and GPT-6 Astra, eight calls each, one per role, single-shot: the brief and packet in the prompt, the review returned as the final message, no tools, no file reads. Reasoning effort xhigh on all three. Single-shot is both cheaper (round one's cost was the agent loop re-reading its context; USD 34 of USD 41 was cache traffic) and fairer, since Codex already runs that way.

**Blind judging.** Per role, a triple labelled A, B, C by a recorded coin (`work/r2/key.json`). Three judges (Opus, Fable, Astra), blind, each scoring every output: the worked direction on the four questions, 0 to 2 each (8); the first step on whether it was done and done correctly, 0 to 2; the missing-analysis on whether it says something about the project the record does not, 0 to 2; the unlimited direction on could-be-wrong and new, 0 to 2; maximum 14. Then a ranking of the three, and five picks across all twenty-four outputs.

**The comparison rule.** Per judge, three pairwise sign tests from the rankings (Fable against Opus, Fable against Astra, Opus against Astra), ties excluded, two-sided p < 0.05. With eight roles a judge finds a difference only at 8 to 0; that is the power this round has, and it is stated here so a 6 to 2 is not later called a finding. "Quality differs" is concluded for a pair only if two of three judges find it in the same direction. Mean scores and mean ranks per lineage per judge are reported as description. Round one's Opus-against-Astra result stands separately; the two rounds are not pooled.

**Cost.** Cap USD 100 in Claude tokens at list rates for the round (the author's figure); the editor's estimate is USD 20 to 30. Codex calls: nine (eight makers, one judge), on the author's plan.

**What would count against this amendment.** The same as the design's, plus: round two's brief steering toward a new genre of its own (the "first step done here" may favour desk derivations over field designs; the judges are asked to score the step on correctness, not on kind); the eight-role power limit being read as evidence of no difference.

**Deviation found at the outside read (20 September 2026).** Both packets carried the claims graph's nodes without its `confidencePolicy`, which prices S1's number on the weak reading only. Makers and judges therefore read 0.8 as the diagnosis's grade. The omission is the editor's; it is recorded here and corrected in `10_directions.md`, not in the outputs.
