Research program / public evidence snapshot

Ask what changed.
Then break it.

Chess is the testbed. We separate state, proposal, selection, search, and value with matched interventions, measurable outcomes, and explicit falsifiers.

Evidence cutoffAug 8, 2026, 8:00 PM UTC
“A score tells us that a system changed. A mechanism test asks which intervention made it change.”

Full games expose external validity (whether findings carry beyond this setup). Paired decisions (the same decision under both conditions) expose causes. The public record keeps those objects separate.

Current boundary. Mechanism headlines remain gated. 7/12 submission gates have passed.

01Design

Defines the intervention, outcome, and failure condition. It is not a model result.

02Preregistered (rules set before outcomes)

Freezes the estimand (the target quantity the analysis is designed to estimate) and decision rules before outcomes. Amendments stay visible.

03Observed

Reports a public aggregate (combined summary) with its sample, uncertainty, and claim boundary.

04Refused

Leaves a headline unavailable when the evidence gate is still open.

Mechanism map

Five ways a move can fail.

Each panel names the stage, the instrument that isolates it, the observation that would count against it, and the strongest current public evidence. Sensitivity (checking whether a result changes when a stated analysis condition is varied) is kept separate from the primary result.

Testable mechanism

Does access to board and legal-action representations change the decision when the chess state is held fixed?

Instrument

A105 matched 2 × 2 factorial: board content × complete legal-action list, equal-length masks, the same full FEN and action universe, and two panel orders.

Falsifier

After transport qualification, a held-out run with no regret contrast would weaken the grounding account; any byte drift outside registered spans invalidates the comparison.

Design · local preflight Current public evidence. A105 is local preflight only; 0 observed model outcomes and no paid execution authorization. Open A105

Featured design instrument

A105 holds the chess state fixed and changes only registered access to redundant board and legal-action representations.

Current evidence

The public snapshot, counted.

Clean games observed320both cohorts · observation-time projection
Frozen Luna curve220protocol-specific games in the curve
Terminal aggregates (final combined results)4E025 · E028 · E029 · A113
Submission gates7/12mechanism headline gated

Reading rule. The counts above describe public artifacts and observation state. They do not turn a terminal aggregate into a general chess-strength or mechanism claim.

Observed results

What survived a completed test.

These rows report public terminal aggregates. The result is narrower than the mechanism label beside it; each boundary is carried into the linked experiment.

Selection

Does candidate order change action choice beyond repeated-call instability?
Adjusted excess 5.2 pp; 95% interval (an uncertainty range calculated by resampling positions) -2.1 pp to 13.5 pp.

INSTRUMENT48 positions · balanced ABBA/BAAB order controls
EVIDENCE STATEObserved result · no clear signalRead E025

Proposal guidance

Does showing a fixed candidate set guide the selected action?
Paired membership effect (difference across matched positions) 93.75 pp; quality remains unmeasured.

INSTRUMENT48 positions · 192 registered calls · zero missing cells (all registered responses present)
EVIDENCE STATEObserved terminal aggregateRead E028

Local consequence

Does exact advice change move utility on the frozen five-piece population?
Paired utility effect (the difference in exact 50-move-rule outcome scores between arms) 7.08 pp; 95% interval 2.92 pp to 11.67 pp.

INSTRUMENT60 positions · 240 calls · 60 complete positions (both arms observed)
EVIDENCE STATEObserved result · qualified positiveRead E029

Engine scaffold

Does showing up to two unranked fixed-budget Stockfish 18 moves change complete-episode survival for Luna?
Paired survival effect (difference across matched arms) 11.60 pp; source-game 95% interval 7.52 pp to 15.99 pp with the allowed missing-outcome check reported separately.

INSTRUMENT320 source puzzles · 640 scheduled episodes · 319 complete pairs
EVIDENCE STATEObserved terminal aggregate · model plus harnessRead A113

Designs and preregistered work

Specified is not observed.

A design can be tested, rendered, and audited before a model runs. Those checks make an experiment ready; they do not supply its outcome.

State access

Does redundant access to the root board or complete legal-move list change system regret when the underlying chess state, action universe, response contract, and request length stay fixed?
Local preflight only; 0 observed model outcomes.

INSTRUMENT2 × 2 factorial · 4 arms · 256 synthetic cells
EVIDENCE STATEPreregistered design · execution gatedRead A105

Registered contract

Which target quantity, missingness rule, and decision rule were fixed before outcomes?
4 registered commitments (rules fixed before outcomes); analyzer lock deviation disclosed (analysis code was completed after collection).

INSTRUMENT60 positions · 240 registered calls · 2 replicates/arm
EVIDENCE STATEPreregistered contract · result separateRead protocol

Evaluation chronology / interim observation

One program.
Four snapshot measures.

Concurrency makes elapsed time, active time, and accumulated suite work answer different questions. This public view is derived from reconciled append-only snapshots and excludes operational and account data.

Observed wall74.4 hours

Elapsed time between the first append-only snapshot and this evidence cutoff.

Active wall74.4 hours

Observed wall time with at least one recognized evaluation suite active.

Suite-worker327.8 hours

The integral of concurrent active suites over observed wall time.

Five-way equivalent65.6 hours

Suite-worker hours divided by five; this is normalized work, not elapsed time.

Evidence boundary. Snapshot intervals provide observation bounds, not exact process launch or exit timestamps. No account usage, credential balance, or model-specific dollar attribution is included. These clocks describe evaluation operations and do not support a strength, causal, or terminal DeepSeek claim.

Cutoff Aug 6, 2026, 6:02 AM UTC · source identity d6b9d1726427

Standards of evidence

Make the falsifier visible.

  1. 01Hold positions and action universes fixed for causal comparisons.
  2. 02Store randomized candidate order (which candidate appeared in each slot) and treatment assignment exactly.
  3. 03Hide scores and ranks when they are part of the intervention.
  4. 04Report provider failures separately; never convert them to losses.
  5. 05Carry N, uncertainty, missingness, and amendments with every estimate.
  6. 06Give every mechanism an explicit observation that could count against it.
Read the evidence doctrine