Two Unranked Engine Moves Helped Luna Survive
In a frozen paired puzzle benchmark, showing Luna two unranked engine-proposed moves improved registered decision survival by 11.60 percentage points. This is evidence for an engine scaffold—not model learning, Elo gain, or general chess strength.
The sharp question was not whether an engine can play chess. It was whether a small engine-generated action scaffold changes what the same language model can do when everything else stays fixed.
A113 paired 320 source-puzzle episodes. In one arm Luna saw only the board. In the other it also saw two engine-proposed legal moves, randomized and unranked. Either arm could still submit any legal move. The registered outcome was the fraction of decisions in the puzzle line that the episode survived.
Of 320 intended pairs, 319 were complete. The paired engine-top-two minus board-only survival effect was +0.115987, or +11.60 percentage points. A 20,000-draw source-game bootstrap gave a 95% interval of [+7.52, +15.99] percentage points. Forty-five discordant pairs succeeded only with the engine scaffold; eight succeeded only board-only. The exact two-sided McNemar p-value was 0.0000002368.
The registered promotion decision passed every gate.
Figure — A113 paired engine-scaffold effect
The missing pair does not decide the sign
One board-only episode was provider-missing; the matched engine-top-two episode was complete. The all-intended-pair sensitivity interval was [+11.56, +11.88] percentage points. Even assigning the missing outcome adversarially therefore does not erase the positive effect.
This matters because “319 complete pairs” and “640 terminal episodes” are not the same claim. All 640 scheduled episodes terminalized, but one ended without an authoritative board-only outcome. The primary estimate correctly uses 319 complete pairs and reports the missingness sensitivity separately.
The scaffold changed both choices and survival
The engine-top-two adoption rate was 0.5294: on just over half of eligible decisions, Luna selected one of the displayed moves. Mean registered decisions survived rose from 0.3323 board-only to 0.6375 with the scaffold. Puzzle-line move accuracy rose descriptively from 0.2663 to 0.4444, and conservative protocol success improved by 0.16875.
Those secondary measures make a purely formatting-based explanation less plausible, but they do not identify the internal mechanism. Luna may have copied a candidate, recognized it as strong, reduced its search space, or deferred to the engine. A113 was designed to establish the intervention effect, not to choose among those explanations.
Where the effect appeared
The paired effect was positive for both agent colors: +9.43 points as Black and +13.75 as White. It was also nonnegative in every registered rating band. The descriptive effects were +30.38 points at 800–1399, +11.25 at 1400–1999, +3.75 at 2000–2599, and +1.25 at 2600+. By episode length they were +18.92 points for two decisions, +5.51 for three, and +4.55 for four.
The gradient is interesting, but it is not a separately powered discovery. Harder and longer puzzles may leave less room for a two-move scaffold, or the band composition may differ in other ways. These slices are descriptive leads for the next preregistered experiment, not independent confirmatory claims.
What improved—and what did not
The agent protocol improved: a fixed Luna model paired with a two-move engine scaffold survived more registered puzzle decisions than the same model with the board alone.
The base model did not learn during this experiment. No weights changed, and A113 does not show improved unassisted performance after the scaffold is removed. It does not produce an Elo estimate, establish stronger full-game play, or prove transfer to other models, positions, engines, prompts, or tools.
This distinction is the research program in miniature. “Agentic chess skill” is not one scalar inside a model. It can emerge from the interaction among a policy model, candidate generator, information boundary, and decision protocol. A controlled scaffold effect is useful precisely because it tells us which component changed.
The next serious benchmark should test transfer rather than merely repeat the headline: held-out puzzle sources, deceptive and partially correct candidate sets, a tool-request policy with an explicit cost, and a post-scaffold board-only evaluation. Those conditions can separate persistent learning from temporary access to a better proposal surface.
Reference destinations are classified from this public copy. Hash-only references return here; unavailable references retain a safe label but expose no private path or mutable artifact. External URLs are not verified by this build.
- Public essay links
- 0
- External URLs
- 0 (not verified)
- In-page hash references
- 0
- No public destination
- 0