Two Exact Moves Changed the Submitted Policy
Showing Luna two unranked tablebase moves improved exact 50-move-rule utility on a frozen synthetic five-piece population. It did not establish Elo gain, learning, or general chess strength.
E029 asked a narrower question than a game ladder can answer. If Luna receives the same board but also sees two unranked moves drawn from an exact Syzygy policy, does its submitted move preserve more 50-move-rule utility?
The terminal result is bounded by a source amendment and a single-source limitation. Source Selection Amendment 1 was accepted after the immutable tablebase labels were known but before any model response; it selected the 60-position v2 corpus from the frozen 360-candidate pool without generating or querying new positions. The amended corpus was qualified against one frozen Syzygy source and adapter, not independently replicated truth. The pre-terminal failure and recovery chronology is preserved in essay 28 (In-page hash only).
Across all 60 registered positions and 240 completed calls, the answer was positive for this frozen cohort. Mean utility was 0.425 in the board-only arm and 0.495833 in the exact-top-two arm. The preregistered within-position effect was +0.070833, or 7.08 percentage points. Its 20,000-draw whole-position bootstrap 95% interval was [0.029167, 0.116667]. The paired sign-flip test gave a two-sided p-value of 0.0032399676.
The registered decision rule classified the result as positive confirmatory: all 60 positions were complete, the mean and bootstrap lower bound were above zero, unresolved-cell bounds were point-identical at +0.070833, and the legality/protocol-success change was +0.016667 rather than an adverse shift.
That classification needs a visible qualification. The estimand, bootstrap draw count and seed, sign-flip rule and seed, missingness bounds, legality threshold, and decision rule were written before collection. The paid execution lock did not bind the analyzer's executable bytes. The first terminal analyzer was incomplete; the deterministic implementation was completed after terminal data under Analysis Completion Amendment 3. The registered rules did not change, but this was not a fully pre-hash-locked analysis implementation.
What the intervention changed
Both arms used the same unrestricted move-string tool. Either arm could submit any legal move. The exact-top-two arm merely displayed two legal moves from the frozen exact tablebase policy, in randomized order and without scores or ranks. The board-only arm did not see them.
Exact-top-two membership rose from 40 of 120 actions (0.333333) board-only to 116 of 120 (0.966667) with advice. The paired membership effect was +0.633333. This establishes that the prompt information changed the submitted policy. The utility result adds that, on this population, the change was beneficial under the preregistered exact 50-move-rule score.
What remained fixed
The population contained 60 synthetic five-piece positions selected before any Luna outcome under Source Selection Amendment 1. It balanced six material families and the side-to-move by root-outcome cells. Each position received two fresh calls per arm in a frozen ABBA or BAAB sequence. Provider failures would have remained missing; none occurred. Malformed or illegal returned actions remained observed zero-utility outcomes.
The source and execution history matters. As essay 28 (In-page hash only) records, the original v1 selection failed closed because a registered stratum was infeasible, with zero Luna calls; a provider-free minimum-cost-flow amendment then selected the feasible v2 corpus from already labelled candidates. During execution, a UCI-only boundary rejected legal SAN before any scored event, a first recovery draft was rejected in preflight because it would have blocked legitimate repeated calls, and Recovery Amendment 2 used attempt identities with the same UCI-or-unambiguous-SAN normalization in both arms.
Those incidents are not decorations around the result. They are part of its identity.
What the result does not say
E029 does not show that Luna learned chess. It does not estimate unassisted skill, full-game win rate, or Elo. It does not separate recognition, copying, search reduction, or deference as internal mechanisms. It does not establish transfer beyond this frozen synthetic five-piece population, this dated Luna snapshot, or this exact prompt-and-tool protocol.
The median paired utility contrast was zero even though the mean was positive. The effect therefore came from a subset of positions rather than a uniform improvement everywhere. Future work should test held-out positions, both colors, misleading or partially correct advice, and whether an agent can learn when to request or distrust candidate information.
Reference destinations are classified from this public copy. Hash-only references return here; unavailable references retain a safe label but expose no private path or mutable artifact. External URLs are not verified by this build.
- Public essay links
- 0
- External URLs
- 0 (not verified)
- In-page hash references
- 2
- No public destination
- 0