Claim observatory / Public claim boundary · August 2026

What we know.
What we do not.

A score becomes science only when its boundary is explicit. This snapshot maps each headline to its evidence status, the inference it cannot support, and the next test that could change our mind.

01Observe

Preserve the surprising result exactly.

02Bound

Name every inference the data cannot identify.

03Falsify

Design the intervention that separates live explanations.

04Replicate

Promote only after held-out evidence survives.

Public evidence registry

Nine claims. Four evidence states. Zero hidden upgrades.

A hand-curated public snapshot. It is not generated from the mutable claims ledger and contains no live operational data. Operational state, private artifacts, and live cross-model conclusions are excluded.

Showing 9 of 9 public claim boundaries.

What the evidence supports

In the frozen Black-only Luna pilot, the joint local rating was 1067 without candidates and 2832 with two randomized, unranked engine candidates.

Claim boundary

This measures the complete protocol, not unaided model strength. The initial candidate sets came from separate time-limited searches and were not guaranteed to be nested.

Decisive next test

Replicate with both colors, fixed-node nested candidates, multiple opening families, and confirmatory samples near each crossing.

Read the linked evidence note