Claim observatory / Public claim boundary · August 2026

What we know.
What we do not.

A score becomes science only when its boundary is explicit. This snapshot maps each headline to its evidence status, the inference it cannot support, and the next test that could change our mind.

01Observe

Preserve the surprising result exactly.

02Bound

Name every inference the data cannot identify.

03Falsify

Design the intervention that separates live explanations.

04Replicate

Promote only after held-out evidence survives.

Public evidence registry

9 claims. Four evidence states. Zero hidden upgrades.

A hand-curated public snapshot through C023. It is not generated from the mutable claims ledger and contains no live operational data. C024-C027 are represented outside this curated observatory by separate sanitized terminal result capsules; omission here is not evidence rejection or publication authorization. Operational state, private artifacts, and live cross-model conclusions are excluded.

Promotion contract / public snapshot

Every claim travels with its boundary.

05
Source identity

Public claim boundary · August 2026 · schema v1. Each row keeps its claim ID and names the public essay carrying the linked evidence note.

Evidence state

Claim disposition and source state remain separate. Frozen, provisional, and protocol/design labels never upgrade the claim text.

Falsifier

The expanded “Falsifier / next test” is the curated next-test field: a public challenge to the boundary, not an outcome already observed.

Omissions

A hand-curated public snapshot through C023. It is not generated from the mutable claims ledger and contains no live operational data. C024-C027 are represented outside this curated observatory by separate sanitized terminal result capsules; omission here is not evidence rejection or publication authorization. The observatory intentionally stops at C023.

Separate terminal capsules

C024-C027 are represented outside this curated observatory by sanitized terminal result capsules on separate routes. Their omission here is not evidence rejection or publication authorization.

Showing 9 of 9 public claim boundaries.

Source identity

C001 · public essay 01
Two Moves Are Better Than Ten
01-two-moves-are-better-than-ten

Frozen evidence
Claim disposition

Frozen observation

What the evidence supports

In the frozen Black-only Luna pilot, the joint local rating was 1067 without candidates and 2832 with two randomized, unranked engine candidates.

Claim boundary

This measures the complete protocol, not unaided model strength. The initial candidate sets came from separate time-limited searches and were not guaranteed to be nested.

Falsifier / next test

Replicate with both colors, fixed-node nested candidates, multiple opening families, and confirmatory samples near each crossing.

Read public source note Permalink to this boundary