# Lab decision log

Format follows `research/judgement/DECISION_LOG.md`: date, decision, level, who decided, alternatives considered, cost and cap, what would reverse it.

## 2026-09-17 — Adopt the lab; first batch of five; cap

**Decision.** The founder adopted the lab under the rules in `00_SEED.md`, approved the recommended first batch (8 Fable's dissent, 1 the case against, 7 the framing study on Fable, 9 four models one question, 5 the sonnet) and the cap (USD 120 build in Claude Opus 5 and Fable tokens, USD 10 API for the test, 3 Codex calls), with the words "Sounds good. Don't worry too much about the tokens. As long as it's reasonable."

**Level.** 2 (adopt, cap) and 3 (batch). Publication of any piece remains his: he reads everything on the branch preview before a merge.

**Alternatives considered.** Tests-and-records only (no made things): rejected for now because the sonnet is the lab's own quality test; if it fails review the lab narrows to that. A different first batch: offered, not taken.

**Cost.** Estimates in `00_SEED.md`. Actuals are logged per piece in each piece's record. The framing study on Fable is priced separately in its pre-registration; Fable's list rates (USD 10 per million input tokens, USD 50 per million output, with thinking billed as output) make the USD 10 API figure in the seed an under-estimate, and the pre-registration states the real estimate and asks for the cap before any real call.

**What would reverse it.** A first-batch piece that passes review and still reads as unserious to a hostile academic reader; the dissent or the case against turning out to be soft; actual cost exceeding the cap by more than a third.

**Held.** 4 (spend), 11 (a named person), 16 (his data), 17 (a child), 18 (money and reputation), 20 (his book).

## 2026-09-17 — Open, not yet decided: who reviews piece 9 from outside the lineage

**The ask.** The lab's standing outside-lineage reviewer is GPT-6 Astra (OpenAI), and piece 9 (`/lab/one-question.html`) reports on GPT-6 Astra and Claude Fable 5.1 and publishes a judgement about Astra's samples. Asking Astra to review it asks a model to certify a record of its own output. A second point sits beside it: the site's outside verdicts are obtained through the Codex wrapper, which this piece sets aside as a route for an answer to a bare question.

**Options.** (a) This piece is read by a model that is neither subject, and Astra keeps the standing role for everything else. (b) Astra reads the protocol and the write-up only, and its verdict goes in the reception ledger with the conflict stated. (c) Both: an outside read from a third model and Astra on the write-up.

**Level.** 3 (options presented, the founder decides). Seed rules 4 and 5.

**Cost.** One Codex call under option (b), or a call to a third provider under (a); within the first batch's three-call allowance.

**Status.** Raised on the piece's Review section, which says the decision is the author's before the piece goes live. Nothing is decided and nothing is spent.

## 2026-09-18 — First batch built; outside reads; cost against the cap

**What happened.** The scaffold and all five pieces were built on branch `site/lab` (checkpoint cef58c3 and after), tested (99 site tests, 314 offline harness tests), and read from outside the lineage by GPT-6 Astra through the Codex command-line tool with no human relay. Verdicts, all filed beside the pieces and served under `downloads/lab/`:

| piece | outside read | state |
|---|---|---|
| 8 Fable's dissent | HOLD, then HOLD narrowing (two items withdrawn, four narrowed, one recentred) | revised to both reads; third read waits on the author |
| 1 the case against | HOLD (two blockers, seven majors) | not revised; the reviewer names five stronger charges the record supports |
| 7 the framing study on Fable | HOLD (nine findings on the filing's accounting, none on the design) | not revised; cannot be filed until it is; cap not asked for |
| 9 one question, several models | ADOPT AMENDED (six exact replacements) | applied the same day |
| 5 the sonnet | HOLD (form holds; three images rhyme-bought; sestet a versified definition) | one revision round in progress; second read pending |

**Cost, actual.** Claude Opus 5 agent work priced at list rates from the agents' own usage records (input USD 5, output USD 25, cache read USD 0.50, cache write USD 6.25 per million tokens): first build workflow USD 191.89 (31 agents), of which roughly USD 65 was work re-run after the desktop app quit twice mid-workflow; first repair workflow USD 35.51; second repair workflow (case findings, runner hardening) USD 12.88; sonnet revision attempts USD 7.82 so far. Total USD 248.10. Provider calls: USD 0.0251 (piece 9) and USD 0.0177 (piece 7 smoke, six draws). Codex calls on the author's plan: six reads (sonnet 116k tokens, dissent 432k and 1.05M, case 1.50M, framing and one question together 1.90M; most of each cached prefix). Fable tokens for drafting the dissent and the pre-registration, adjudication and these records are not metered here.

**Against the cap.** The approved cap was USD 120 for the build. The build stands at about USD 248 at list rates (including the second repair workflow, USD 12.88, and two sonnet revision attempts that stalled on output length, USD 7.82), about twice the cap: roughly a quarter of the overrun is the duplicated work after the app quits, the rest is the review-and-fix cycles costing more than the seed's per-piece estimates (the case against alone ran to about USD 90 across drafting, merging, review, one blocked fix and one completed fix). The seed's own clause, that actual cost exceeding the cap by more than a third counts against the costing method, fires. **Consequence taken:** no further Opus agent rounds are started without the author's decision; the remaining repairs (the case against, the framing filing, a third dissent read) are proposals with their own caps, below.

**Level.** 1 for the work above (executed, author notified here). 2 for what follows.

**Proposals for the author, each with a cap.**
1. Revise the case against to the outside read (the reviewer's five stronger charges are specific and the quotations already check out): one Opus round with recheck, cap USD 25, then one Codex read.
2. Revise the framing filing to the outside read (documentation only; no provider calls): cap USD 15 in Opus or done by Fable; then the run's own cap, USD 15, if the author wants the run.
3. Third read of the dissent: one Codex call, no Opus.
4. The sonnet: if the second read holds, withdraw it and keep the attempt listed as withdrawn (no cost); if it adopts, go live with the batch.
5. Go-live for whatever is adopted: changelog line, feed, reception ledger entries (five, one per piece, under the ledger's convention), merge. Reputational note: every piece carries its reads on the page, including the HOLDs, which is the lab's design and also what a hostile reader will see first.

**What would reverse this.** The author deciding that the lab's cost per adopted piece is not worth it; the record then stays on the branch, unmerged, as a complete account of a first batch that was built, read, and mostly held.

## 2026-09-18 (later) — The sonnet withdrawn after its second outside read

**Decision.** GPT-6 Astra's second read of the revised sonnet returned HOLD, narrowing to one major on the sestet. The lab had said before the read that a second hold would withdraw the piece; it is withdrawn, kept on its page and on the index as withdrawn, with both verdicts served beside it. Level 1 (executed as pre-announced; the author can reverse it by funding a third revision: Fable by hand plus one Codex read, no Opus).

**Alternatives considered.** A third revision round now: rejected because the rule was announced in advance and each round spends the author's Codex plan. Removing the page: rejected because the lab keeps what it withdraws.

**What it says about the lab.** One made thing, two reads, no adoption. The seed's clause that a lab which cannot make things well should keep to tests and records is now live as a question for the author, not settled.

## 2026-09-18 (evening) — The author's decision on the batch: finish the records, repair the filing, leave the sonnet withdrawn

**Decision.** The author approved the package proposed after the reads, with the words "Go for it": (1) the dissent goes for a third outside read as it stands; the case against gets one more Opus round against the reviewer's named charges, then one read; (2) Fable repairs the framing filing by hand, with the USD 15 run still a separate decision once the filing is ready; (3) the sonnet stays withdrawn; (4) whatever adopts goes live in one commit; (5) nothing merges before that. Cap: USD 50 in Opus, three Codex reads, Fable's time. Level 2.

**Alternatives offered.** Only (1) and (4); or stop and keep the branch as the record. Not taken.

**Cost of the approved package so far (18 September, evening).** Case against, round 2 (one Opus revision at xhigh plus one recheck): USD 24.50 at list rates, inside its USD 25 cap. Framing filing and dissent repairs: Fable by hand, no Opus. Codex reads on the author's plan: dissent third read 1.71M input tokens (1.61M cached), case second read and framing second read pending. Running Opus total for the lab: about USD 272.60.

## 2026-09-18 (night) — Go-live of the first batch

**Decision.** Executed under the author's approval of the post-read package: the lab goes live on main with the one-question piece, the dissent and the case against adopted (amended), the framing filing adopted as a filing and not run, and the sonnet withdrawn and kept. Five reception entries (14 to 18), counts 18/18/14 under the ledger's stated convention, a dated changelog entry, the feed regenerated, the working contract's programme map updated. Level 2, approval naming the consequence, the artefacts and the amount.

**Codex reads in the package.** Dissent third read 1.71M input tokens; case second read 1.96M; framing second read 2.30M (all mostly cached prefix), on the author's plan.

**Open.** The framing run's USD 15 abort threshold (the author's); reconciling the reception ledger's earlier nine with its convention (put to the ledger by the dissent and the case); the sonnet, reopenable.

## 2026-09-18 (late) — Author's decisions: copy pass merged; framing run approved at USD 15; sonnet retried on two more lineages; the floor seeded

**Decisions (the author, "Go for it! And yes on the two decisions").** (1) The plain-English copy pass by GPT-6 Astra is merged to main. (2) The framing study on Claude Fable 5.1 runs under the filed pre-registration with `--max-usd 15` as an abort threshold; the filing is committed before the run and its hash goes in the manifest. (3) The sonnet is retried with Claude Fable 5.1 agents and with GPT-6 Astra agents under the same six constraints and two-judge protocol, so three lineages can be compared; stated caps: USD 60 in Fable tokens, about twelve Codex calls on the author's plan, no Opus. (4) Fable writes the seed for the live research floor and the conjecture type as level-2 proposals; nothing runs from them without his read.

**Level.** 2 throughout, approvals naming consequence, artefact and amount.

## 2026-09-19 — The sonnet on three lineages, judged blind

**What happened.** The sonnet brief was run again on Claude Fable 5.1 agents (six attempts, two judges, one revision: USD 21.80 at Fable list rates against a USD 60 cap) and on GPT-6 Astra agents through Codex (nine calls on the author's plan). The three winners were then judged blind by one judge from each lineage (Opus USD 0.93, Fable USD 1.08, Astra one Codex call). All three judges ranked the Fable poem first; the Astra judge, the one reader outside its lineage, took it with a light revision at two lines, recorded and not applied. Record `05d_sonnet_three_lineages.md`, page `site/lab/sonnet-three.html`, branch `lab/cycle-2`, not merged.

**Level.** 1 for the runs (within the caps the author set, "Go for it"); 2 for going live, which is the author's.

**What it says about the lab.** The first made thing to pass an outside-lineage read as the kind of thing it is. The standard still names a human poetry editor, whom none of the three has met.

## 2026-09-19 — The author's "Let's proceed": cycle 2 goes live; the floor and the conjecture type adopted; PR 84 reconciled

**Decisions (the author, "Let's proceed").** (1) `lab/cycle-2` merges to main: the three-lineage sonnet comparison (adopted at its blind cross-lineage read), the framing study's no-primary result, and reception entry 19. (2) The floor seed (`research/floor/00_SEED.md`) and lab amendment A1 (the conjecture type) are adopted as proposed; the first cycle runs at its own cap (USD 40 in Claude tokens, three Codex calls); the monthly cap is not yet set by the author and the seed's USD 150 stands as the proposal until he names a figure. (3) GPT-6 Astra's draft PR 84 (editorial lab index) is reconciled with cycle 2 by Fable and merged with the amendments in Fable's review. (4) A second framing run under a preamble-tolerant parse rule is not started; it waits on a separate yes.

**Level.** 2 throughout.

## 2026-09-19 — Floor cycle 1 complete: two conjectures adopted with amendments, one held

**What happened.** Under the floor seed's first-cycle cap (USD 40, three Codex calls), three conjectures were proposed, one per lineage, and each read from outside its lineage on amendment A1's four questions. Adopted with amendments: what "as it seems" keeps (GPT-6 Astra, read by Claude Opus 5, seven replacements applied) and a strong reading of S1 without P2 (Claude Fable 5.1, read by GPT-6 Astra, eight replacements applied). Held: a persistence criterion without a metric (Claude Opus 5, read by GPT-6 Astra). Actual cost USD 15.70 in Claude tokens plus three Codex calls. Record: `research/floor/cycle-1/`, branch `floor/cycle-1`, nothing merged, nothing live. Level 1 for the runs; level 2 for what follows.

**Open for the author.** Whether the adopted conjectures become lab pieces and go live; one revision round for the held one (about USD 5) or drop; the monthly cap; cycle 2.

## 2026-09-19 — Floor cycle 2 complete; the floor's first outputs go live

**What happened.** Under the author's "run with the rest as you see fit" and the first cycle's cap, cycle 2 proposed a revised persistence conjecture (held a second time; parked), a duration-container conjecture (held; the packet omitted the Study B items, the editor's fault) and the same-model-ids record (adopted with nine amendments). Cost: USD 13.92 in Opus tokens for the makers, about USD 5 for the agent applying the record's amendments, six Codex calls of which three were near-empty (launched before the files were on the branch; the editor's error). Cycle 1's two adopted conjectures and cycle 2's record go live in the lab in one merge with reception entries 20 to 25 (counts 25/25/18), a changelog entry and the feed. Level 2, under the author's standing approval of the package.

**Open for the author.** The persistence line's third round or parking; the duration-container revision; the monthly cap; cycle 3's openings.

## 2026-09-19 — The floor's monthly cap set: USD 150

**Decision (the author, "Agree on the cap").** The floor runs under a monthly cap of USD 150 in Claude tokens at list rates, per-test provider calls capped at USD 20, four Codex calls per cycle, as the seed proposed. September's cycles 1 and 2 (about USD 34.62 in Claude tokens) count against it. Level 2.

## 2026-09-19 — Process finding: the plain-words explanation of the parked line did not carry

**What happened.** Asked what the parked persistence line was about, the editor answered in short plain sentences using the line's own terms (order, metric, occurrences, recurrence). The author: "The plain words sound important but they don't carry meaning to me. I want to understand what you're trying to say." Under the working contract this is a process finding (rule 5: if he has to say "make it clear", log it and fix the rule).

**Rule fixed.** Any line put to the author for a decision is explained by analogy first (one everyday image carrying the structure), each term mapped to the image, the outside reads' objections stated in the image's terms, then the stakes for the wager, then the options. Test before sending: he could say the problem back from the message alone. Recorded in the editor's memory as well.

**Level.** 0 (a rule on the editor's own conduct). Cost: none.

## 2026-09-19 — Floor cycle 3 opens: the persistence line's third round, adopted with amendments; the round's cap overrun

**Decision (the author, "Only $10? ... Let's go!! Run with it").** One more narrow round on the persistence conjecture, capped at USD 10 in Claude tokens plus one Codex call, scoped to the two blockers of the second hold.

**What happened.** Claude Opus 5 revised (three commitments, an open question, a counting rule, a rejection condition per commitment); two Opus refuters and an auditor returned fifteen findings before the outside read, all applied; GPT-6 Astra's third read returned ADOPT AMENDED with five exact replacements, applied by the editor, after which the critic's stated verdict is ADOPT. Record: `research/floor/cycle-3/`, branch `floor/cycle-3`, page and surfaces prepared, nothing merged. Level 1 for the round; level 2 for going live.

**Cost and the overrun.** USD 16.70 in Claude Opus 5 tokens at list rates against the USD 10 cap, plus one Codex call. The overrun is the editor's: the cap was stated for one reviser and one read, and the editor added three refuting agents and a fix pass without restating the cap. Rule adopted (level 0, the editor's conduct): a change of scope after a cap is stated goes back to the author as a new cap or does not happen. September's floor spend against the USD 150 monthly cap: cycles 1 and 2 USD 34.62, this round USD 16.70, total USD 51.32; the filing agent that built the page and surfaces ran on Claude Sonnet 5 (about 216,000 tokens, about USD 1 at list rates), bringing September's floor spend to about USD 52.

**Alternatives considered.** Park the line after two holds (the seed's default), rejected by the author for one narrow round. Spend the Codex call without in-lineage refutation, rejected by the editor because the previous two reads had each found definitional gaps a cheaper reader could have caught; the outcome supports that, and the cost rule above is the price.

**Open for the author.** Going live (level 2); the rest of cycle 3.

## 2026-09-19 — The persistence conjecture goes live; an agent may carry an invitation on stated terms

**Decisions (the author, "I'm good with 1 and 2").** (1) `floor/cycle-3` merged to main at 14155ef: the persistence conjecture adopted at its third read is live at `site/lab/conjecture-persistence.html`; reception 26/26/19. Level 2. (2) The floor seed's human-collaborator rule 1 is amended: an agent may carry an invitation only when the author has approved the exact text, the first line discloses that an AI agent wrote and sent it for the project and names the author as responsible, it goes from a project address the author owns and can read, and the reply is treated as the person chooses; approval is per message. Level 3, taken by the author with the reputational risk stated first (a recipient may publish the message as an example of machine mail; that risk is priced, not removed).

**Not decided.** Any actual send. The author: "Actual outreach needs to wait until I read more about Kammerer." A reading note from the record is at `research/floor/outreach/00_kammerer_reading.md`. The project mailbox (level 4) is the author's to create; options are in the same note.

**Alternatives considered.** Keep the rule as it was, with the author sending himself: he declined ("I'm not going to do it"). A standing scope for the agent (send within a list without per-message approval): rejected by the editor as too wide for a first contact under his name.
