Experiment A113 / terminal content-addressed aggregate (final combined result identified by its exact file contents)

Two moves.
A local scaffold effect.

A113 attributes the result to a model-plus-harness protocol (model plus execution setup): the same Luna model saw up to two unranked moves (shown without scores or ranks) from a fixed-budget Stockfish 18 proposer, while either arm could submit any legal move.

320 source puzzles640 scheduled episodes319 complete pairs (both arms observed for one source puzzle)Luna model + fixed-budget Stockfish 18 proposer harness
Registered questionDoes showing up to two unranked Stockfish moves change complete-episode survival?
Pairing unit320 source-puzzle episodes crossed by arm

319 pairs are complete for the primary estimate.

Primary outcomefraction of registered puzzle decisions survived

The estimate is paired survival (the difference in complete-episode survival for matched arms), not a rating or learning measure.

AttributionModel plus harness

Luna model + fixed-budget Stockfish 18 proposer harness; the route does not isolate a model-only capability.

Primary paired result / registered gates passed

+11.60 pp

Engine-top-two minus board-only complete-episode survival was +11.60 pp across 319 complete pairs from 320 source puzzles and 640 scheduled episodes. The source-game bootstrap 95% interval (a resampling-based uncertainty range over source games) was +7.52 pp to +15.99 pp.

McNemar's exact two-sided p-value (a matched-pair test of arm differences) was 0.0000002368 (45 scaffold-only successes versus 8 board-only successes).

Promoted
A113 paired engine-scaffold effect: the scaffold improved complete-episode survival by 11.60 percentage points, with a 95 percent bootstrap interval from 7.52 to 15.99 points.
The route-specific A113 figure shows the paired survival estimate and its source-game bootstrap interval. It is an aggregate visual; puzzles, moves, prompts, responses, and provider identifiers remain outside the public surface.

Missingness sensitivity (checking how the result changes when the missing pair takes allowed extreme outcomes)

The missing pair does not decide the sign.

The public capsule records 1 provider-missing pair (one matched pair without a provider result). The all-intended-pair sensitivity interval (range under the allowed missing-outcome extremes) is +11.56 pp to +11.88 pp; the sign remains robust. The missing outcome is not silently converted into a loss or folded into the 319-pair primary denominator.

Terminal boundary. 640 episodes were scheduled, but the primary paired result uses only the 319 complete pairs and reports missingness separately.

Board-only protocol success15.36%

Descriptive registered protocol success.

Engine-top-two protocol success32.50%

Descriptive registered protocol success.

Scaffold adoption52.94%

Eligible decisions selecting a displayed move.

Discordant pairs (matched pairs with different outcomes)53

45 scaffold-only · 8 board-only.

Claim boundary

Protocol effect.
Not model learning.

What the artifact supports. A positive paired engine-scaffold effect on registered puzzle-decision survival for this Luna cohort; no learning, Elo, unassisted-strength, or general agentic-chess claim.

  • This estimates the local effect of an engine-generated two-move scaffold on a frozen paired puzzle protocol.
  • It does not show that model weights changed, that the base model learned, or that unassisted chess skill improved.
  • It is not an Elo estimate and does not establish transfer to full games, other models, or other tool protocols.
  • Subgroup effects are descriptive and are not separately powered confirmatory claims.
  • The experiment does not distinguish copying, recognition, search reduction, or deference as mechanisms.
Open C027 evidence boundary

Publication boundary. This terminal evidence capsule carries no publication authorization for a broader publication claim, model-only capability claim, Elo estimate, or general transfer claim.

Public evidence identity

The page consumes the content-addressed terminal aggregate and its route-specific deterministic figure. It does not recompute the estimate or expose the underlying puzzle corpus, moves, prompts, responses, tokens, costs, account state, or operational paths.