Claim observatory / Public claim boundary · August 2026
What we know. What we do not.
A score becomes science only when its boundary is explicit. This snapshot maps each headline to its evidence status, the inference it cannot support, and the next test that could change our mind.
01Observe
Preserve the surprising result exactly.
02Bound
Name every inference the data cannot identify.
03Falsify
Design the intervention that separates live explanations.
04Replicate
Promote only after held-out evidence survives.
Public evidence registry
Nine claims. Four evidence states. Zero hidden upgrades.
A hand-curated public snapshot. It is not generated from the mutable claims ledger and contains no live operational data. Operational state, private artifacts, and live cross-model conclusions are excluded.
Showing 9 of 9 public claim boundaries.
What the evidence supports
In the frozen Black-only Luna pilot, the joint local rating was 1067 without candidates and 2832 with two randomized, unranked engine candidates.
Claim boundary
This measures the complete protocol, not unaided model strength. The initial candidate sets came from separate time-limited searches and were not guaranteed to be nested.
Decisive next test
Replicate with both colors, fixed-node nested candidates, multiple opening families, and confirmatory samples near each crossing.