Public experiment registry / evidence before headline
Questions first.
Results with boundaries.
Each route starts with the question, names the public evidence state, and leads to the result the artifact can support—or the reason no model result exists yet.
5 public routes4 terminal aggregates1 design / preflight
Read the claim contract →01QuestionStart with the intervention and the decision it is meant to isolate.
02StateSeparate a design or preflight from an aggregate result.
03ResultFollow only the public estimate, interval, and disposition.
04BoundaryKeep the unanswered inference visible beside the result.
Index / public artifacts
A question-to-result
reading path.
01The cards keep design, public readout, and limitation in the same frame. Open a route when you want the instrument, the aggregate, and its fuller provenance.
- Question
- Does redundant access to the root board or complete legal-move list change system regret when the underlying chess state, action universe, response contract, and request length stay fixed?
- Public readout
- 0 observed model outcomes and 0 provider calls are admitted. The public readout is a version-bound local preflight, not a model run.
- Evidence boundary
- R3's installed Responses path remains the observed NO-GO; R3A is isolated and not integrated into the live harness. No provider-received bytes, provider compatibility, provider behavior, model effect, chess-quality effect, power result, final-position result, paid execution, or causal benefit has been observed.
- Question
- Does changing candidate order add instability beyond repeated identical-order calls?
- Public readout
- Adjusted excess disagreement was +5.2 pp; the 95% interval was −2.1 pp to +13.5 pp. The interval includes zero, so excess order sensitivity was not clearly supported.
- Evidence boundary
- Action stability for one dated Luna alias cohort on development-exposed Black positions; not move quality, Elo, search, or a universal null.
- Question
- Does showing one fixed five-move suggestion set change Luna's selected action?
- Public readout
- Visible fixed-set membership was 95.8% versus 2.1% in board-only; the paired effect was +93.8 pp with a 95% interval of +88.5 pp to +97.9 pp.
- Evidence boundary
- Development-only causal effect of showing one fixed five-move suggestion set on Luna set adherence; no move-quality, Elo, vision, search, or candidate-generation claim.
- Question
- Do two exact tablebase moves improve local 50-move-rule utility?
- Public readout
- Exact-top-two utility was 49.6% versus 42.5% board-only; the paired effect was +7.1 pp with a whole-position 95% interval of +2.9 pp to +11.7 pp. The registered decision is positive confirmatory; confirmatory statistics implement prospectively registered rules, but the executable analyzer was completed after terminal data and was not hash-bound in the execution lock
- Evidence boundary
- A local causal effect of displaying two unranked exact-policy moves on exact 50-move-rule move utility for this cohort; no full-game transfer, Elo, learning, vision, search, or general chess-strength claim.
- Question
- Does showing up to two unranked fixed-budget Stockfish 18 moves change complete-episode survival for Luna on frozen source puzzles?
- Public readout
- The paired engine-scaffold survival difference was +11.6 pp with a source-game 95% interval of +7.5 pp to +16.0 pp. This is a local protocol effect for 319 complete pairs, not a model-only learning result.
- Evidence boundary
- A positive paired engine-scaffold effect on registered puzzle-decision survival for this Luna cohort; no learning, Elo, unassisted-strength, or general agentic-chess claim.