Research program / August 2026 → February 2027
From leaderboard
to mechanism.
Our north star is not the strongest chess harness. It is a reproducible chain from observation to mechanism to intervention to held-out replication.
“A scalar score can tell us that a system changed. It cannot tell us which part learned to see.”
Chess is the measurement environment. The objects of study are broader: grounded state, proposal quality, comparative selection, search, value estimation, uncertainty, and long-horizon coherence.
Elo remains the external-validity test. Paired decision experiments are the main instrument for causal understanding.
Mechanism map
Five failures hiding inside one move.
Latent failure stage
Can the agent reconstruct the board it is acting on?
Matched FEN, board, piece-list, and deliberately inconsistent interfaces.
If representation changes legality but not move quality, grounding was not the main bottleneck.
Featured causal instrument
A105 holds the chess state fixed and randomizes access to redundant board and legal-action representations.
Portfolio
Six linked workstreams.
External validity
Can the complete agent survive the consequences of its own decisions?
Proposal
Does a strong move enter the model’s consideration set?
Selection
Can the model choose the best action already available?
State
Does the interface preserve a grounded, legal board model?
Search + value
Can verified consequences extend reliable decision depth?
Efficiency
What strength is purchased per token, second, tool call, and dollar?
Critical path
Six months, gated by evidence.
Freeze the current generation
Terminal cohorts, content-addressed evidence, no mutable dependency.
Resolve the retrospective signal
Separate proposal regret from selector regret without causal overclaim.
Build the paired substrate
600 locked positions, nested candidates, independent labels, grouped splits.
Identify the mechanism
Candidate-width replication, overload controls, notation and order ablations.
Search, value, architecture
Promote modules one at a time on held-out decisions.
Replicate and release
Benchmark v1, model cards, correction log, paper, clean-room reproduction.
Standards of evidence
The result must be harder to publish than to imagine.
- 01Identical positions for causal comparisons.
- 02Randomized order, stored exactly.
- 03Scores and ranks hidden unless intervened on.
- 04Provider failures reported, never converted to losses.
- 05Every estimate travels with N and uncertainty.
- 06Every mechanism owns an explicit falsifier.