Claim observatory / Public claim boundary · August 2026
What we know. What we do not.
A score becomes science only when its boundary is explicit. This snapshot maps each headline to its evidence status, the inference it cannot support, and the next test that could change our mind.
01Observe
Preserve the surprising result exactly.
02Bound
Name every inference the data cannot identify.
03Falsify
Design the intervention that separates live explanations.
04Replicate
Promote only after held-out evidence survives.
Public evidence registry
9 claims. Four evidence states. Zero hidden upgrades.
A hand-curated public snapshot through C023. It is not generated from the mutable claims ledger and contains no live operational data. C024-C027 are represented outside this curated observatory by separate sanitized terminal result capsules; omission here is not evidence rejection or publication authorization. Operational state, private artifacts, and live cross-model conclusions are excluded.
Promotion contract / public snapshot
Every claim travels with its boundary.
05
Source identity
Public claim boundary · August 2026 · schema v1. Each row keeps its claim ID and names the public essay carrying the linked evidence note.
Evidence state
Claim disposition and source state remain separate. Frozen, provisional, and protocol/design labels never upgrade the claim text.
Falsifier
The expanded “Falsifier / next test” is the curated next-test field: a public challenge to the boundary, not an outcome already observed.
Omissions
A hand-curated public snapshot through C023. It is not generated from the mutable claims ledger and contains no live operational data. C024-C027 are represented outside this curated observatory by separate sanitized terminal result capsules; omission here is not evidence rejection or publication authorization. The observatory intentionally stops at C023.
Separate terminal capsules
C024-C027 are represented outside this curated observatory by sanitized terminal result capsules on separate routes. Their omission here is not evidence rejection or publication authorization.
Showing 9 of 9 public claim boundaries.
Source identity
C001 · public essay 01 Two Moves Are Better Than Ten 01-two-moves-are-better-than-ten
Frozen evidence
Claim disposition
Frozen observation
What the evidence supports
In the frozen Black-only Luna pilot, the joint local rating was 1067 without candidates and 2832 with two randomized, unranked engine candidates.
Claim boundary
This measures the complete protocol, not unaided model strength. The initial candidate sets came from separate time-limited searches and were not guaranteed to be nested.
Falsifier / next test
Replicate with both colors, fixed-node nested candidates, multiple opening families, and confirmatory samples near each crossing.
C015 · public essay 08 Building a Chess-Native Reasoning Architecture 08-building-a-chess-native-reasoning-architecture
Protocol / design
Claim disposition
Frozen no-signal
What the evidence supports
A102's frozen Luna replay found zero attempts matching the preregistered sequence: a production-reachable legal command, an explicit revision cue, and a different unambiguous legal move.
Claim boundary
This is a content-addressed no-signal result in the retained corpus. It does not establish first-command precedence, final-answer precedence, or the prevalence of revisions in new traffic.
Falsifier / next test
Keep the production precedence gate unresolved until a held-out natural corpus or prospective natural/semisynthetic protocol supplies identifying revision cases.
C016 · public essay 08 Building a Chess-Native Reasoning Architecture 08-building-a-chess-native-reasoning-architecture
Protocol / design
Claim disposition
Frozen observation
What the evidence supports
A110A's locked immediate-prompt audit found 11,002 of 11,002 direct all-legal selections inside the displayed candidate set; a separate post-hoc same-position sensitivity found 23,332 of 23,353 inside.
Claim boundary
The 100% locked-primary membership and 99.91% post-hoc sensitivity are parser-membership observations, not selector quality, move value, causal benefit, or Elo.
Falsifier / next test
Run the paired proposal-versus-selection decomposition with independent engine labels and source-game clustered inference.
C021 · public essay 08 Building a Chess-Native Reasoning Architecture 08-building-a-chess-native-reasoning-architecture
Protocol / design
Claim disposition
Preregistered test
What the evidence supports
A prospective instrument will reveal aligned or misaligned successor states, opponent replies, and short continuations while holding candidate proposals fixed.
Claim boundary
A balanced provider-free design audit is not model evidence and cannot establish a calculation defect or scaffold benefit.
Falsifier / next test
Execute the registered paired cells only after the position pool, independent evaluator, prompt audit, and separate budget gates pass.
C022 · public essay 16 The Tool Call That Looked Like Causality 16-the-tool-call-that-looked-like-causality
Frozen evidence
Claim disposition
Frozen observation
What the evidence supports
The frozen chronology shows accepted legal commands frequently followed board or legal-list queries in the same position chain.
Claim boundary
Tool choice, returned prompt, attempt depth, prior failure, candidate condition, and position difficulty were entangled. Chronology is not information use or causal benefit.
Falsifier / next test
Randomize byte-matched board and legal-list panels on identical frozen positions, with tool access disabled during the measured decision.
C023 · public essay 21 The Patch That Did Not Erase the Failure 21-the-patch-that-did-not-erase-the-failure
Protocol / design
Claim disposition
Preregistered test
What the evidence supports
R3 observed the installed Responses adapter rewrite required submit_move choice to auto in 128 pre-network captures. R3A separately wrapped a fresh synthetic client, changed only that normalized tool_choice field, and preserved required submit_move across 128 captures and 32 arm-invariant blocks.
Claim boundary
R3 remains the immutable primary installed-stack NO-GO. R3A is a version-bound local remediation candidate for AutoGen 0.11.2, OpenAI Python 2.30.0, and httpx 0.28.1; it is not integrated into the live chess harness, proves neither provider compatibility nor provider-received bytes, does not authorize paid A105 execution, and contains no model, final-position, move-quality, power, or causal outcome.
Falsifier / next test
Keep the installed Responses path blocked. Independently review and integrate the candidate only in a non-production harness, repeat the full capture, then require real model/route binding and a separate provider-compatibility or received-request gate before any paid run.