Open research program · August 2026
The harness is
part of the player.
We use chess to separate what language-model agents can represent, propose, compare, calculate, and value—then build experiments that can prove each explanation wrong.
Frozen Luna pilot · 220 complete games
More options did not mean more strength.
Candidate widths came from separate time-limited MultiPV searches. This frozen curve is an observation, not a causal bandwidth effect. The paired, nested replication is preregistered.
From score to mechanism
A leaderboard says how much.
A laboratory asks why.
A loss can hide five different failures. We give each one its own intervention, outcome, and falsifier—because different bottlenecks demand different systems.
Latent failure stage
Can the agent reconstruct the board it is acting on?
Matched FEN, board, piece-list, and deliberately inconsistent interfaces.
If representation changes legality but not move quality, grounding was not the main bottleneck.
Evidence architecture
Observation → mechanism → intervention → replication.
Every public headline must resolve to frozen inputs, executable analysis, machine-readable exclusions, uncertainty, and an explicit claim boundary.
Open the Claim Observatory →Laboratory notebook
Latest writing
Laboratory operations
The Analysis That Existed Before the Results
A causal experiment is not preregistered if its most consequential decisions are still waiting inside the analysis script…
Laboratory operations
The Control Arm Was Full of Chess
Before paying for 4,800 model answers, we made the intervention prove that it actually existed…
Laboratory operations
The Tool Call That Looked Like Causality
A model asked for the board, then made a move. The hard part is proving what the board changed.…