Research program / public evidence snapshot
Ask what changed.
Then break it.
Chess is the testbed. We separate state, proposal, selection, search, and value with matched interventions, measurable outcomes, and explicit falsifiers.
“A score tells us that a system changed. A mechanism test asks which intervention made it change.”
Full games expose external validity (whether findings carry beyond this setup). Paired decisions (the same decision under both conditions) expose causes. The public record keeps those objects separate.
Current boundary. Mechanism headlines remain gated. 7/12 submission gates have passed.
Defines the intervention, outcome, and failure condition. It is not a model result.
Freezes the estimand (the target quantity the analysis is designed to estimate) and decision rules before outcomes. Amendments stay visible.
Reports a public aggregate (combined summary) with its sample, uncertainty, and claim boundary.
Leaves a headline unavailable when the evidence gate is still open.
Mechanism map
Five ways a move can fail.
Each panel names the stage, the instrument that isolates it, the observation that would count against it, and the strongest current public evidence. Sensitivity (checking whether a result changes when a stated analysis condition is varied) is kept separate from the primary result.
Testable mechanism
Does access to board and legal-action representations change the decision when the chess state is held fixed?
A105 matched 2 × 2 factorial: board content × complete legal-action list, equal-length masks, the same full FEN and action universe, and two panel orders.
After transport qualification, a held-out run with no regret contrast would weaken the grounding account; any byte drift outside registered spans invalidates the comparison.
Design · local preflight Current public evidence. A105 is local preflight only; 0 observed model outcomes and no paid execution authorization. Open A105
Featured design instrument
A105 holds the chess state fixed and changes only registered access to redundant board and legal-action representations.
Current evidence
The public snapshot, counted.
Reading rule. The counts above describe public artifacts and observation state. They do not turn a terminal aggregate into a general chess-strength or mechanism claim.
Observed results
What survived a completed test.
These rows report public terminal aggregates. The result is narrower than the mechanism label beside it; each boundary is carried into the linked experiment.
Selection
Does candidate order change action choice beyond repeated-call instability?
Adjusted excess 5.2 pp; 95% interval (an uncertainty range calculated by resampling positions) -2.1 pp to 13.5 pp.
Proposal guidance
Does showing a fixed candidate set guide the selected action?
Paired membership effect (difference across matched positions) 93.75 pp; quality remains unmeasured.
Local consequence
Does exact advice change move utility on the frozen five-piece population?
Paired utility effect (the difference in exact 50-move-rule outcome scores between arms) 7.08 pp; 95% interval 2.92 pp to 11.67 pp.
Engine scaffold
Does showing up to two unranked fixed-budget Stockfish 18 moves change complete-episode survival for Luna?
Paired survival effect (difference across matched arms) 11.60 pp; source-game 95% interval 7.52 pp to 15.99 pp with the allowed missing-outcome check reported separately.
Designs and preregistered work
Specified is not observed.
A design can be tested, rendered, and audited before a model runs. Those checks make an experiment ready; they do not supply its outcome.
State access
Does redundant access to the root board or complete legal-move list change system regret when the underlying chess state, action universe, response contract, and request length stay fixed?
Local preflight only; 0 observed model outcomes.
Registered contract
Which target quantity, missingness rule, and decision rule were fixed before outcomes?
4 registered commitments (rules fixed before outcomes); analyzer lock deviation disclosed (analysis code was completed after collection).
Evaluation chronology / interim observation
One program.
Four snapshot measures.
Concurrency makes elapsed time, active time, and accumulated suite work answer different questions. This public view is derived from reconciled append-only snapshots and excludes operational and account data.
Elapsed time between the first append-only snapshot and this evidence cutoff.
Observed wall time with at least one recognized evaluation suite active.
The integral of concurrent active suites over observed wall time.
Suite-worker hours divided by five; this is normalized work, not elapsed time.
Evidence boundary. Snapshot intervals provide observation bounds, not exact process launch or exit timestamps. No account usage, credential balance, or model-specific dollar attribution is included. These clocks describe evaluation operations and do not support a strength, causal, or terminal DeepSeek claim.
Cutoff Aug 6, 2026, 6:02 AM UTC · source identity d6b9d1726427…Standards of evidence
Make the falsifier visible.
- 01Hold positions and action universes fixed for causal comparisons.
- 02Store randomized candidate order (which candidate appeared in each slot) and treatment assignment exactly.
- 03Hide scores and ranks when they are part of the intervention.
- 04Report provider failures separately; never convert them to losses.
- 05Carry N, uncertainty, missingness, and amendments with every estimate.
- 06Give every mechanism an explicit observation that could count against it.