Public experiment registry / evidence before headline

Questions first.
Results with boundaries.

Each route starts with the question, names the public evidence state, and leads to the result the artifact can support—or the reason no model result exists yet.

5 public routes4 terminal aggregates1 design / preflight
Read the claim contract
01Question

Start with the intervention and the decision it is meant to isolate.

02State

Separate a design or preflight from an aggregate result.

03Result

Follow only the public estimate, interval, and disposition.

04Boundary

Keep the unanswered inference visible beside the result.

Index / public artifacts

A question-to-result
reading path.

01

The cards keep design, public readout, and limitation in the same frame. Open a route when you want the instrument, the aggregate, and its fuller provenance.

  1. 01

    A105 / Protocol design

    Matched root-information factorial ablation

    Design · local preflight only
    Question
    Does redundant access to the root board or complete legal-move list change system regret when the underlying chess state, action universe, response contract, and request length stay fixed?
    Public readout
    0 observed model outcomes and 0 provider calls are admitted. The public readout is a version-bound local preflight, not a model run.
    Evidence boundary
    R3's installed Responses path remains the observed NO-GO; R3A is isolated and not integrated into the live harness. No provider-received bytes, provider compatibility, provider behavior, model effect, chess-quality effect, power result, final-position result, paid execution, or causal benefit has been observed.
    128 rendered cells · 256 synthetic analysis cellsOpen A105 design
  2. 02

    E025 / Terminal aggregate

    Candidate-order sensitivity

    Terminal · not clearly supported
    Question
    Does changing candidate order add instability beyond repeated identical-order calls?
    Public readout
    Adjusted excess disagreement was +5.2 pp; the 95% interval was −2.1 pp to +13.5 pp. The interval includes zero, so excess order sensitivity was not clearly supported.
    Evidence boundary
    Action stability for one dated Luna alias cohort on development-exposed Black positions; not move quality, Elo, search, or a universal null.
    v1.1 corrected primary · 48 positions · 192 callsOpen E025 result
  3. 03

    E028 / Terminal paired aggregate

    Prompt guidance

    Terminal · fixed-set guidance effect
    Question
    Does showing one fixed five-move suggestion set change Luna's selected action?
    Public readout
    Visible fixed-set membership was 95.8% versus 2.1% in board-only; the paired effect was +93.8 pp with a 95% interval of +88.5 pp to +97.9 pp.
    Evidence boundary
    Development-only causal effect of showing one fixed five-move suggestion set on Luna set adherence; no move-quality, Elo, vision, search, or candidate-generation claim.
    48 paired positions · 192 registered calls · 0 provider-error cellsOpen E028 result
  4. 04

    E029 / Terminal aggregate

    Exact advice protocol

    Terminal · positive result / analysis deviation
    Question
    Do two exact tablebase moves improve local 50-move-rule utility?
    Public readout
    Exact-top-two utility was 49.6% versus 42.5% board-only; the paired effect was +7.1 pp with a whole-position 95% interval of +2.9 pp to +11.7 pp. The registered decision is positive confirmatory; confirmatory statistics implement prospectively registered rules, but the executable analyzer was completed after terminal data and was not hash-bound in the execution lock
    Evidence boundary
    A local causal effect of displaying two unranked exact-policy moves on exact 50-move-rule move utility for this cohort; no full-game transfer, Elo, learning, vision, search, or general chess-strength claim.
    60 positions · 60/60 complete · post-run analyzer disclosedOpen E029 result
  5. 05

    A113 / Terminal content-addressed aggregate

    Engine scaffold survival

    Terminal · model-plus-harness effect
    Question
    Does showing up to two unranked fixed-budget Stockfish 18 moves change complete-episode survival for Luna on frozen source puzzles?
    Public readout
    The paired engine-scaffold survival difference was +11.6 pp with a source-game 95% interval of +7.5 pp to +16.0 pp. This is a local protocol effect for 319 complete pairs, not a model-only learning result.
    Evidence boundary
    A positive paired engine-scaffold effect on registered puzzle-decision survival for this Luna cohort; no learning, Elo, unassisted-strength, or general agentic-chess claim.
    320 source puzzles · 640 scheduled episodes · 319 complete pairs · 1 missing pairOpen A113 result