bootstrap interval for the paired effect
Public evidence / idle snapshot / Aug 8, 2026, 8:00 PM UTC
Evidence before
explanation.
A public laboratory for understanding model-plus-harness decision systems in chess. Every result carries its protocol, snapshot, and claim boundary.
Public terminal artifact / E029 Terminal aggregate
+7.08 pp local utility effect.
Exact-advice effect on 50-move-rule move utility for one Luna snapshot and one frozen synthetic five-piece population. Across 60/60 complete paired positions, the cohort completed 240 of 240 registered calls. No provider-missing calls.
Boundary. A local causal effect of displaying two unranked exact-policy moves on exact 50-move-rule move utility for this cohort; no full-game transfer, Elo, learning, vision, search, or general chess-strength claim. Qualification. confirmatory statistics implement prospectively registered rules, but the executable analyzer was completed after terminal data and was not hash-bound in the execution lock
mean exact utility
mean exact utility with advice
the positive mean was not uniform
Secondary / terminal content-addressed scaffold effect / publication withheld
A113.
Still secondary.
319/320 complete pairs; +11.60 pp paired survival difference.
Boundary. A local model-plus-harness puzzle result—not a broader publication claim, model-only capability claim, Elo estimate, or general transfer claim.
Open the A113 terminal aggregate →Evidence states / Aug 8, 2026, 8:00 PM UTC
One snapshot.
Four labels.
Observed names the current ledger snapshot. Public terminal artifacts are terminal aggregates. Planned names registered targets. Gates keep claims bounded.
Open the status ledger →Observed / non-causal
A curve is a clue.
Not a cause.
The 220-game Luna curve describes 5 complete candidate conditions under a Black-only protocol against Stockfish 18 with UCI LimitStrength.
Protocol boundary. These are local protocol-specific measurements, not FIDE, online-platform, or general model ratings. Candidate conditions are complete agent protocols, not isolated estimates of base-model chess strength. Candidate sets are not guaranteed nested, so the curve does not estimate a causal candidate-count effect.
Inspect the protocol leaderboard →Three reader paths