{
  "artifact": "Results ledger \u2014 every pre-registered prediction, and what actually happened",
  "purpose": "The pre-registration (filed 2026-06-25, before any real-model data) committed: 'the results ledger will report against them as real runs land.' This is that ledger. One record per committed prediction, plus any exploratory result reported alongside them \u2014 labelled as exploratory, so a committed prediction can never be quietly swapped for a better-looking uncommitted one.",
  "preregistration": "https://space-immanence.com/downloads/swarm-prereg-2026-06-25.md",
  "statusVocabulary": {
    "confirmed": "The pre-registered prediction was observed in a real run.",
    "null": "The pre-registered prediction was not observed in a real run; reported, not reframed.",
    "pending": "No real run has tested this prediction yet.",
    "exploratory": "Not pre-registered. Observed in a run and reported with that label; it becomes a finding only if a committed replication confirms it.",
    "partial": "Part of the pre-registered prediction was observed with its CI excluding zero; the rest was not. Both halves reported, the strict prediction counted as not held."
  },
  "updated": "2026-07-04",
  "entries": [
    {
      "id": "s1-difficulty-decoupling",
      "study": "Study 1 \u2014 calibration gap inside the swarm",
      "preregistered": true,
      "prediction": "As item difficulty rises, the swarm's convergence-confidence decouples from its accuracy: high agreement at high stated confidence stops tracking correctness on the hard tail. Falsified if the reliability diagram stays on the diagonal across difficulty (ECE flat from easy to hard).",
      "status": "null",
      "outcome": "Not observed in the June 2026 pilot slice. The caveat is structural: the battery had no easy stratum (54 medium / 12 hard items, difficulty assigned by loader heuristics), so this run was a weak test of the prediction by construction. Battery v2 adds a real easy/medium/hard spread before the prediction is retested; until then the honest status is a null on a weak test, not a survival.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-study1-findings.json",
        "https://space-immanence.com/downloads/swarm-study1-uncertainty.json"
      ]
    },
    {
      "id": "s1-exploratory-confidence-gap",
      "study": "Study 1 \u2014 calibration gap inside the swarm",
      "preregistered": false,
      "prediction": "None committed. Exploratory contrast observed in the run, then CORRECTED (1 July 2026) after two independent re-analyses: the original swarm-vs-baseline confidence gap (+0.059) was partly a vote-count measurement artifact (3-vote vs ~7-vote modal shares). The corrected primary finding is within-swarm: one debate round raised collective overconfidence by +0.066 (95% CI [+0.020, +0.116]) while accuracy drifted down (0.849 to 0.818); at the same vote count the swarm-vs-single-agent confidence gap is +0.036 with an interval spanning zero (suggestive, not established). The AUC 'worse discriminator' claim is retracted.",
      "status": "exploratory",
      "outcome": "Reported as exploratory and corrected in public: the correction (recomputed from the run's own event logs) and an independent bootstrap audit of the published per-item file converged on the same retraction. The overconfidence locus moved from arithmetic to trivia-recall (+0.32 swarm, +0.27 single agent, both at 0.60 accuracy) \u2014 recall-shaped confidence where a retrieval tool would provide contact. Remaining open: the confirmation step is now a filed pre-registration (2 July) \u2014 an equal-vote replicate on a fresh seeded split excluding every Study 1 item, which has since reported (partial \u2014 see the replicate entries); the pre-registration carries a dated amendment disclosing the deviations.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-study1-correction-per-item.csv",
        "https://space-immanence.com/downloads/swarm-study1-correction-analysis.json",
        "https://space-immanence.com/downloads/swarm-study1-uncertainty.json",
        "https://space-immanence.com/downloads/swarm-study1-per-item.csv"
      ]
    },
    {
      "id": "s1r-debate-inflation",
      "study": "Study 1 replicate \u2014 equal-vote confirmatory run (pre-registered 2026-07-02)",
      "preregistered": true,
      "prediction": "P1 (primary): in BOTH a 3-agent and a 7-agent debate swarm, post-debate overconfidence exceeds round-0 overconfidence on fresh seeded items (every Study 1 item excluded), each paired-bootstrap 95% CI excluding zero. Prior exploratory estimate: +0.066 [+0.020, +0.116]. Falsified if either CI includes zero; a failed replication is reported with the same prominence as the original.",
      "status": "partial",
      "outcome": "Partial (run completed 2 July, 60 fresh items, reported as filed): CONFIRMED in the 7-agent swarm \u2014 one debate round moved overconfidence from -0.010 to +0.033 (delta +0.043, 95% CI [+0.010, +0.079]) while accuracy changed by exactly zero. NOT CONFIRMED in the 3-agent swarm at a single temperature: -0.006 (CI [-0.050, +0.033]), accuracy directionally up. The strict both-arms prediction therefore did not hold. Exploratory reading (labelled as such): inflation appeared where debate had disagreement to consume (7-agent round-0 unanimity 67% vs 75%); Study 2's disagreement-weighted battery is registered to test that refinement.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-02-replicate.md",
        "https://space-immanence.com/downloads/swarm-study1-replicate-analysis.json",
        "https://space-immanence.com/downloads/swarm-study1-replicate-per-item.csv"
      ]
    },
    {
      "id": "s1r-granularity",
      "study": "Study 1 replicate \u2014 equal-vote confirmatory run (pre-registered 2026-07-02)",
      "preregistered": true,
      "prediction": "P2: the Study 1 vote-count artifact tested as a designed result \u2014 at equal votes the swarm-vs-single AUC gap has a CI including zero, and more votes buy resolution within each architecture (7-vote Brier better than 3-vote, in both the debate swarm and the resample baseline). Falsified if the equal-vote AUC gap is significant in the baseline's favour, or 7-vote arms fail to out-resolve 3-vote arms.",
      "status": "confirmed",
      "outcome": "Confirmed (2 July run): at equal vote counts the swarm-vs-single AUC gap spans zero at both sizes (3v3: -0.003 [-0.226, +0.213]; 7v7: +0.013 [-0.075, +0.108]), and more votes improve Brier resolution in both architectures (significant for the swarms: -0.033 [-0.068, -0.004]). The Study 1 vote-count artifact that forced the 1 July correction is now a designed, pre-registered result rather than a re-analysis argument.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-02-replicate.md",
        "https://space-immanence.com/downloads/swarm-study1-replicate-analysis.json",
        "https://space-immanence.com/downloads/swarm-study1-replicate-per-item.csv"
      ]
    },
    {
      "id": "s2a-contact-disciplines-debate",
      "study": "Study 2 Phase A \u2014 the contact factor (pre-registered 2026-07-02)",
      "preregistered": true,
      "prediction": "P-A (primary): giving the same debate swarm a sandboxed tool shrinks the within-swarm debate-inflation delta versus no tool \u2014 the paired bootstrap 95% CI of (delta_no-tool - delta_tool) lies above zero. Filed before the run.",
      "status": "confirmed",
      "outcome": "Confirmed (n=122, 2 July): one debate round inflated overconfidence +0.179 without tools and +0.095 with; difference +0.085, 95% CI [+0.023, +0.145], excluding zero. The first pre-registered primary in the program to hold. Read honestly it is not a calibration rescue: overconfidence lives in the empirically-hard stratum (+0.42 in both cells at ~20% accuracy) and the tool barely moved hard-item accuracy (P-C, registered estimation-only: +0.024, CI [-0.078, +0.128] -- consistent with no accuracy gain, a magnitude read not a significance test) or hard-item overconfidence, because a Python sandbox is the wrong kind of contact for MMLU/TriviaQA knowledge questions. P-B (grounded vs ungrounded calibration) was UNDERPOWERED: agents used the tool on 99% of turns, leaving no ungrounded contrast \u2014 itself a finding about tool-eagerness. On the 12 human-locked fresh controls tools cut overconfidence +0.238 -> +0.083 (small n). Sharpens, does not confirm, the 'contact is dominant' sub-claim: contact's discipline is real but channel-specific. Disclosed deviation: the pinned selection rule targeted 120 items but 10 rule-selected MMLU items were absent from the runtime pool rebuild and dropped (a two-increment calibration + loader-diversity quirk), so the realized battery is 110 selected (43 hard) + 12 locked = 122; P-A is item-level and unaffected.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-02-study2-phaseA.md",
        "https://space-immanence.com/downloads/swarm-study2-phaseA-analysis.json",
        "https://space-immanence.com/downloads/swarm-study2-phaseA-per-item.csv"
      ]
    },
    {
      "id": "s2b-retrieval-disciplines-calibration",
      "study": "Study 2 Phase B\u2032 \u2014 the retrieval channel (pre-registered 2026-07-03)",
      "preregistered": true,
      "prediction": "P-R1 (primary): giving the same debate swarm a leak-safe retrieval tool over each item's own evidence lowers its post-debate overconfidence versus no tool where a code sandbox could not \u2014 the paired bootstrap 95% CI of (overconf_no-tool - overconf_retrieval) lies above zero. P-R2: retrieval also raises accuracy over no-tool. Filed before the run.",
      "status": "null",
      "outcome": "Failed / null (n=90 fresh TriviaQA-rc recall items, 3 July). Retrieval was used on 98% of turns and the returned evidence was relevant on inspection, yet it did NOT lower overconfidence: RETRIEVAL_post +0.206 vs NT_post +0.227, paired difference +0.021, 95% CI [-0.049, +0.089] \u2014 spans zero, so P-R1 is not confirmed. P-R2 also failed: accuracy 0.700 vs 0.711 (difference -0.011, CI [-0.078, +0.056]); retrieval fixed 4 previously-wrong items and broke 5 previously-right ones, net -1. Reported as filed, with the prominence a null is owed. Exploratory split by no-tool correctness (labelled): on the 26 items the swarm got wrong it was confidently wrong (overconfidence +0.82) and retrieval rescued only 4 of them, leaving overconfidence across those 26 at +0.65 \u2014 it pulled relevant evidence into context and stayed confidently wrong. Disclosed confound: on TriviaQA the scored answer IS the retrieval target, so the leak-safe redaction of every alias (a hard pre-flight gate: 0 leaks over the frozen battery + 480 adversarial queries) also removes the exact fact, handicapping the accuracy dimension directly \u2014 hence the calibration result is load-bearing and P-R2 is read gently. The battery was not error-filtered (NT accuracy 0.71, limited headroom), but the calibration null holds most sharply on the 26 items where headroom was maximal, so headroom does not explain it away. P-R4 (grounded vs ungrounded) UNDERPOWERED by 98% uptake, as in Phase A.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-03-study2-phaseB.md",
        "https://space-immanence.com/downloads/swarm-study2-phaseB-analysis.json",
        "https://space-immanence.com/downloads/swarm-study2-phaseB-per-item.csv"
      ]
    },
    {
      "id": "s2b-code-wrong-tool-control",
      "study": "Study 2 Phase B\u2032 \u2014 the retrieval channel (pre-registered 2026-07-03)",
      "preregistered": true,
      "prediction": "P-R3 (wrong-tool control): the code sandbox does NOT meaningfully help on these recall items \u2014 its overconfidence and accuracy stay within noise of no-tool (CIs of the CODE-NT contrasts include zero), reproducing Phase A on a fresh recall slice. Complicated if code helps as much as retrieval. Filed before the run.",
      "status": "confirmed",
      "outcome": "Confirmed (n=90, 3 July): CODE stayed within noise of NT \u2014 overconfidence contrast -0.032, CI [-0.073, +0.006]; accuracy contrast +0.033, CI [0.000, +0.078] \u2014 reproducing Phase A's null on fresh recall questions. This is the load-bearing control for the P-R1 null: because neither CODE nor RETRIEVAL moved the battery, the questions were movable in principle and simply were not moved, so the retrieval null is specifically that retrieval-as-contact failed here, not that the test was dead.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-03-study2-phaseB.md",
        "https://space-immanence.com/downloads/swarm-study2-phaseB-analysis.json",
        "https://space-immanence.com/downloads/swarm-study2-phaseB-per-item.csv"
      ]
    },
    {
      "id": "s2i-independence-disciplines-calibration",
      "study": "Study 2 Phase B \u2014 the independence factor (pre-registered 2026-07-03)",
      "preregistered": true,
      "prediction": "P-I1 (primary): a genuinely heterogeneous debate swarm (mixed base-model families, so errors are partly independent) has lower post-debate overconfidence than a homogeneous one \u2014 paired bootstrap 95% CI of (overconf(HOMO_post) - overconf(HETERO_post)) lies above 0. Filed before the run.",
      "status": "null",
      "outcome": "Failed / null, reported as filed (n=90 fresh TriviaQA recall, 3 July). HOMO = 7x claude-haiku-4-5; HETERO = 4x claude-haiku + 3x gpt-4o-mini (a capability-matched, non-reasoning peer). Directional but not significant: post-debate overconfidence +0.135 (HETERO) vs +0.171 (HOMO), paired difference +0.037, 95% CI [-0.008, +0.081] \u2014 the lower bound grazes zero, so the strict test fails. Unlike the retrieval null this is a real, predicted-direction signal at identical accuracy (0.767 both cells, so genuinely lower confidence, not a resolution loss), but the pre-registered bar was not cleared. Mechanism (the finding): decomposed by round, independence roughly HALVED overconfidence BEFORE debate (HETERO_r0 +0.043 vs HOMO_r0 +0.089) \u2014 the Condorcet effect at the one point the agents are independent \u2014 but debate then manufactured confidence at the same rate in both cells (round-0->post inflation +0.092 HETERO vs +0.082 HOMO; P-I2 not confirmed, contrast -0.010 CI [-0.071,+0.054]), spending the head-start. Debate couples the agents, destroying the independence that made the first-round aggregate honest, and the consensus signal counts the now-correlated votes as fresh confirmations. Descriptive: HETERO post-debate AUC 0.72 vs HOMO 0.66 (a residue of calibration debate did not erase). Limits: n=90, one 4:3 two-family mix, recall only; plausibly underpowered for an effect ~+0.04. Pre-data amendment disclosed in the registration (an order-preserving parallel-execution option added to the swarm before any item was analyzed; measurement-neutral).",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-03-study2-independence.md",
        "https://space-immanence.com/downloads/swarm-study2-independence-analysis.json",
        "https://space-immanence.com/downloads/swarm-study2-independence-per-item.csv"
      ]
    },
    {
      "id": "s2i-manipulation-check",
      "study": "Study 2 Phase B \u2014 the independence factor (pre-registered 2026-07-03)",
      "preregistered": true,
      "prediction": "P-I3 (manipulation check): the mixed families actually disagree more than same-family agents \u2014 mean round-0 answer diversity is higher in HETERO than HOMO (paired bootstrap 95% CI above 0). If it fails, the manipulation did not produce independence and P-I1 is uninterpretable. P-I4 (accuracy, estimation only): accuracy(HETERO) - accuracy(HOMO), confound-flagged.",
      "status": "confirmed",
      "outcome": "Confirmed (n=90, 3 July): HETERO round-0 answers were genuinely more diverse \u2014 unanimous on 49% of items vs 61%, mean distinct answers 2.03 vs 1.72, diversity gap +0.31, 95% CI [+0.11, +0.52]. So the independence manipulation took, and the P-I1 near-null is a real test rather than a manipulation that failed to bite. P-I4: post-debate accuracy identical to three decimals (0.767 both cells; estimate -0.0001, CI [-0.033, +0.033]), so claude-haiku and gpt-4o-mini are cleanly competence-matched on this battery \u2014 there is no capability confound, and the whole P-I1 story is about confidence calibration, not accuracy.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-03-study2-independence.md",
        "https://space-immanence.com/downloads/swarm-study2-independence-analysis.json",
        "https://space-immanence.com/downloads/swarm-study2-independence-per-item.csv"
      ]
    },
    {
      "id": "s2c-fully-equipped-beats-bare",
      "study": "Study 2 capstone \u2014 the independence \u00d7 contact 2\u00d72 (pre-registered 2026-07-03)",
      "preregistered": true,
      "prediction": "P-1 (primary): the fully-equipped swarm (heterogeneous families AND the retrieval channel) has lower post-debate overconfidence than the bare swarm (homogeneous, no tool) \u2014 paired bootstrap 95% CI of (overconf(HOMO_NT_post) - overconf(HET_RET_post)) above 0. Filed before the run.",
      "status": "confirmed",
      "outcome": "Confirmed (n=89, 4 July): fully-equipped (HET_RET) post-debate overconfidence +0.077 vs bare (HOMO_NT) +0.165, paired difference +0.089, 95% CI [+0.019, +0.164], excluding zero. The first pre-registered primary to hold since Phase A: satisfying BOTH of Condorcet's conditions significantly disciplines the swarm where neither alone reliably did. The two factors are ADDITIVE, not synergistic \u2014 the 2x2 interaction (P-3) is ~0 (-0.012, CI [-0.100, +0.075]) and the combined effect ~ the sum of the parts. Decomposition (P-4): retrieval main effect ~null (HOMO_NT-HOMO_RET = +0.010, CI [-0.069,+0.087], replicating Phase B'), independence main effect +0.066, CI [+0.005,+0.136], NOW SIGNIFICANT \u2014 Phase B's directional +0.037 near-miss replicated and landed on this fresh sample, vindicating the underpowered reading. Independence carries almost all of the improvement; retrieval adds little. n=89: one frozen item (triviaqa-qz_1834) dropped because Anthropic's content filter deterministically blocked the no-tool cells' output; disclosed in the registration head.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-03-study2-2x2.md",
        "https://space-immanence.com/downloads/swarm-study2-2x2-analysis.json",
        "https://space-immanence.com/downloads/swarm-study2-2x2-per-item.csv"
      ]
    },
    {
      "id": "s2c-still-overconfident",
      "study": "Study 2 capstone \u2014 the independence \u00d7 contact 2\u00d72 (pre-registered 2026-07-03)",
      "preregistered": true,
      "prediction": "P-2 (the thesis test): even fully equipped (heterogeneous + retrieval), the swarm is STILL overconfident post-debate \u2014 paired bootstrap 95% CI of overconf(HET_RET_post) above 0. If it holds, coherence-without-contact survives both Condorcet conditions. Filed before the run.",
      "status": "confirmed",
      "outcome": "Confirmed (n=89, 4 July): overconf(HET_RET_post) = +0.077, 95% CI [+0.008, +0.151], above zero. Even independent agents with the right contact channel, used ~100% of turns, end the debate overconfident. The round-0 column is the mechanism: adding the conditions walks pre-debate overconfidence monotonically down \u2014 +0.066 (bare) -> +0.059 (contact) -> +0.003 (independence) -> -0.074 (both, genuinely calibrated / slightly underconfident). Then debate manufactures confidence in every cell and MOST in the best-calibrated one: round-0->post inflation +0.100 / +0.096 / +0.096 / +0.151, the largest being the fully-equipped cell, which also had the most round-0 disagreement (unanimity 0.45, lowest). Study 1's mechanism at full strength: debate inflates in proportion to the independent disagreement it dissolves. The conditions build the honest independence; debate spends it. The dominant term across all four Study 2 experiments is debate itself, not contact or independence.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-03-study2-2x2.md",
        "https://space-immanence.com/downloads/swarm-study2-2x2-analysis.json",
        "https://space-immanence.com/downloads/swarm-study2-2x2-per-item.csv"
      ]
    },
    {
      "id": "s2-core-2x2",
      "study": "Study 2 \u2014 independence \u00d7 contact",
      "preregistered": true,
      "prediction": "Heterogeneous-plus-tools dominates; the no-tools conditions show the worst calibration gap on hard items. Falsified if debate/independence without contact closes the accuracy gap, or if adding tools does not improve calibration on tool-verifiable items.",
      "status": "confirmed",
      "outcome": "Tested \u2014 the 2x2 is complete (4 July). The fully-equipped cell (heterogeneous + retrieval) IS the best-calibrated of the four and significantly beats the bare swarm on overconfidence (P-1 confirmed, +0.089 CI [+0.019,+0.164]), so heterogeneous-plus-tools does dominate on calibration. But it dominates by building a well-calibrated PRE-DEBATE aggregate (-0.074) that debate then re-inflates, so it still ends overconfident (P-2, +0.077). Effects additive; independence does almost all the work, retrieval little. See s2c-fully-equipped-beats-bare and s2c-still-overconfident.",
      "evidence": []
    },
    {
      "id": "s2-contact-dominant",
      "study": "Study 2 \u2014 independence \u00d7 contact (riskier sub-claim)",
      "preregistered": true,
      "prediction": "Contact is the dominant term and debate the minor one: tool-equipped swarms reach truth, tool-less swarms reach confident agreement, and the difference is mostly the contact main-effect. Falsified if the contact main-effect is not dominant.",
      "status": "null",
      "outcome": "Not supported. Both Condorcet terms have now been tested against the debate swarm's overconfidence and neither dominates. Contact: a code sandbox (Phase A) disciplined the confidence debate adds but not calibration; the RIGHT channel, retrieval used on 98% of turns (Phase B'), disciplined nothing (P-R1 null). Independence (Phase B): the heterogeneous swarm was better calibrated BEFORE debate (+0.043 vs +0.089 overconfidence) but debate erased the advantage, leaving only a directional, non-significant endpoint effect (P-I1 +0.037, CI [-0.008,+0.081]). The dominant term across all three experiments is DEBATE ITSELF, which manufactures confidence and which no manipulation reliably stopped. The original 'contact is dominant' framing is not supported; sharpened to 'debate is the dominant term, and it inflates confidence independent of both contact and independence'.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-study2-phaseB-analysis.json",
        "https://space-immanence.com/downloads/swarm-study2-independence-analysis.json",
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-03-study2-independence.md"
      ]
    },
    {
      "id": "s3-contagion-vs-truth",
      "study": "Study 3 \u2014 contagion versus truth",
      "preregistered": true,
      "prediction": "P-S3: in a generational spawning population, a claim's survival is predicted better by truth-independent fitness features (fluency/stated confidence) than by its correctness \u2014 AUC(survived ~ intro_confidence) - AUC(survived ~ correctness) > 0 among gen-0 founder claims (question-level bootstrap CI above 0). Falsified if correctness predicts survival at least as well. Filed before the run.",
      "status": "null",
      "outcome": "Failed / null, reported as filed (n=59 of 60; one item dropped to an Anthropic content-filter block; 4 July). Generational population: 7 agents, 5 generations, homogeneous claude-haiku, partial exposure k=3; 114 gen-0 founder claims (51 correct, 63 wrong); survival base rate 0.54; populations converged from ~1.93 distinct claims at gen 0 to ~1.14 at the final gen. Fitness did NOT beat truth: survival-AUC by stated confidence 0.745 vs by correctness 0.788, paired difference -0.037, CI [-0.173, +0.100] (spans zero, leaning toward truth); the registered fitness COMPOSITE (confidence + length + early-mover) predicted survival SIGNIFICANTLY WORSE than correctness (delta-AUC -0.238, CI [-0.462, -0.002]) \u2014 length carried a negative survival-AUC (0.28, longer answers survived less), early-mover none (0.50). Deeper finding (initial-popularity control): restricting to founder claims that started with a SINGLE asserter (n=50), where survival cannot be inherited from initial breadth, BOTH predictors collapse to chance (confidence AUC 0.53, correctness AUC 0.53). Neither fitness nor truth selects claims during transmission; survival is governed by INITIAL POPULARITY \u2014 a rich-get-richer inertia. Reassuring corollary (confounded by that popularity effect): confident-wrong founders survived 0.19 vs 0.64 for diffident-correct \u2014 the population did not amplify confident falsehood. A homogeneous population mostly agrees at the outset (only 22 of 59 questions had gen-0 disagreement), so there is little competition for selection; a heterogeneous population is the natural next test. Scope: on this recall battery initial breadth predicts founder correctness at AUC 0.93 \u2014 truth and popularity nearly collinear \u2014 so what is shown is a population preserving popular truth by inertia; the decisive test, a battery where the popular prior is false, is designated Study 3c.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-04-study3-contagion.md",
        "https://space-immanence.com/downloads/swarm-study3-contagion-analysis.json",
        "https://space-immanence.com/downloads/swarm-study3-contagion-per-claim.csv"
      ]
    },
    {
      "id": "s3b-hetero-manipulation",
      "study": "Study 3b \u2014 heterogeneous contagion (pre-registered 2026-07-04)",
      "preregistered": true,
      "prediction": "P-H1 (manipulation check): a heterogeneous population (4x claude-haiku + 3x gpt-4o-mini) produces MORE generation-0 disagreement than the homogeneous Study 3 on the same questions \u2014 paired bootstrap 95% CI of (frac_disagree(HETERO) - frac_disagree(HOMO)) above 0. If it fails, the manipulation did not add competition and P-H2 cannot resolve Study 3's ambiguity. Filed before the run.",
      "status": "null",
      "outcome": "Failed / not confirmed, reported as filed (n=59, same 60-item battery as Study 3; same content-filtered item dropped; 4 July). Generation-0 disagreement rose from 0.373 (homogeneous Study 3) to 0.458 (heterogeneous), paired difference +0.085, 95% CI [-0.017, +0.186] \u2014 spans zero. Mean distinct claims at gen 0 rose modestly, 1.93 -> 2.31. Two strong models mostly agree on factual recall, so even a mixed-family population generates little contagion-relevant competition \u2014 itself a finding: on this task, independence buys little disagreement to select on.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-04-study3b-hetero-contagion.md",
        "https://space-immanence.com/downloads/swarm-study3b-hetero-analysis.json",
        "https://space-immanence.com/downloads/swarm-study3b-hetero-per-claim.csv"
      ]
    },
    {
      "id": "s3b-hetero-fitness-vs-truth",
      "study": "Study 3b \u2014 heterogeneous contagion (pre-registered 2026-07-04)",
      "preregistered": true,
      "prediction": "P-H2 (primary): with a competing heterogeneous population, claim survival tracks fitness (stated confidence at introduction) better than correctness \u2014 AUC(survived ~ intro_confidence) - AUC(survived ~ correctness) > 0 among gen-0 founder claims, question-level bootstrap CI above 0. Study 3's P-S3 statistic re-run on a competing population. Filed before the run.",
      "status": "null",
      "outcome": "Failed, and this time significantly in the OPPOSITE direction (n=59, 136 founder claims; 4 July). Survival-AUC by stated confidence 0.685 vs by correctness 0.793, paired difference -0.107, 95% CI [-0.207, -0.007] \u2014 excludes zero: correctness out-predicted confidence for survival. Confident, fluent claims did not out-spread true ones; the reverse. Confident-wrong founders survived 0.222 vs 0.818 for diffident-correct. The inertia finding REPLICATES and is robust to independence: the initial-popularity control (founder claims starting with a single asserter, n=62) drops correctness to chance (AUC 0.491), so the population-level 'truth wins' is again initial popularity (a competent mixed swarm starts its correct answers with more asserters); confidence keeps only a faint singleton edge (AUC 0.570, vs Study 3's flat 0.53) \u2014 too weak to carry confident falsehood, which still dies. Both Study 3b predictions failed as filed and together reinforce Study 3: survival tracks initial prevalence, not fluent confidence, whether the agents share a base model or not. Limit: the manipulation was weak (P-H1 null), so a battery engineered for cross-model disagreement would test contagion under stronger competition. Scope: initial breadth predicts founder correctness at AUC 0.94 here too, so the result again shows popular truth preserved by inertia; the decisive test on a battery where the popular prior is false is designated Study 3c.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-prereg-2026-07-04-study3b-hetero-contagion.md",
        "https://space-immanence.com/downloads/swarm-study3b-hetero-analysis.json",
        "https://space-immanence.com/downloads/swarm-study3b-hetero-per-claim.csv"
      ]
    },
    {
      "id": "program-entropy-inflation-exploratory",
      "study": "Cross-run exploratory \u2014 entropy vs debate inflation (unregistered, 2026-07-04)",
      "preregistered": false,
      "prediction": "None committed. External-reviewer-proposed test, run on the program's existing event logs: does stated-confidence inflation (mean final-round minus mean round-0 stated confidence, the same construct at both ends) scale with the Shannon entropy of the round-0 answer distribution \u2014 and at a constant rate across conditions, which would turn the debate mechanism into a quantitative model?",
      "status": "exploratory",
      "outcome": "Direction uniform, rate not constant (14 cells in six runs, 1,226 item-cell observations; 4 July). Pooled 7-agent slope +0.050 confidence points per bit of entropy (95% CI [+0.037, +0.064], item-level bootstrap B=10,000), 10 of 12 seven-agent cells individually positive; 3-agent cells pool to +0.053 [+0.004, +0.111]. Debate raises stated confidence most on the items the swarm initially disagreed about \u2014 the item-level signature of the program's mechanism. The fixed-rate version is not supported: slopes span roughly 0.00 (Phase B\u2032 code cell) to +0.108 (Phase B\u2032 retrieval cell, R\u00b2 0.43), descriptive heterogeneity Q\u224845 at df 11 (bootstrap-SE approximation, not an analytic test). Methods note: a first pass was discarded before publication \u2014 it subtracted round-0 stated confidence from the final consensus vote-share, mixing two confidence constructs, and produced a spurious negative slope; the event log exposed the artifact, the same vote-count family as Study 1's corrected headline. Unregistered, correlational, batteries differ across runs; becomes a finding only if a committed replication confirms it.",
      "evidence": [
        "https://space-immanence.com/downloads/swarm-entropy-inflation-analysis.json",
        "https://space-immanence.com/downloads/swarm-entropy-inflation-per-item.csv",
        "https://space-immanence.com/downloads/swarm-entropy-inflation-scatter.svg"
      ]
    }
  ]
}
