11 frozen essays
A bounded observation tied to frozen public evidence; it is not automatically causal or general.
Public notebook · 30 essays · evidence snapshot
Choose an entry point that matches your question, or follow the numbered sequence. Every essay remains available once, with its current public evidence state attached.
Snapshot chronology / public boundary
Essay numbers preserve research order, not publication dates. The state badge is the public evidence boundary for the essay; it does not upgrade a design or interim note because a later result exists.
Three clocks for one evaluation program. Snapshot intervals provide observation bounds, not exact process launch or exit timestamps. No account usage, credential balance, or model-specific dollar attribution is included. These clocks describe evaluation operations and do not support a strength, causal, or terminal DeepSeek claim.
Cutoff August 6, 2026 at 6:02 AM UTC · source identity d6b9d1726427…Evidence key / 30 essays
A bounded observation tied to frozen public evidence; it is not automatically causal or general.
A dated or interim view; it is not a terminal result.
A protocol, audit, or prospective method; it carries no outcome by itself.
Reading paths / no duplicate links
These paths are thematic entry points through the same numbered chronology. Each public essay appears in one path only; every link keeps its canonical slug.
Path 01 / Measurement
Begin with the surprising result, then attach the full-game protocol before reading a rating as a property of a model. 4 essays · 4 frozen.
With no candidate help, Luna struggled against a 1320-rated Stockfish opponent. Given ten strong moves to consider, it still struggled. Given only two, the combined system held its own against Stockfish at 2820
A chess puzzle begins at the moment the puzzle setter finds interesting. The pieces are already arranged. The tactic exists. The task ends as soon as the right move is named.
Ask for a chess player's Elo and the question sounds ordinary. Ask for a language model's chess Elo and almost every noun becomes ambiguous.
More information should help.
Path 02 / Diagnosis
Trace candidate assistance, interface behavior, and the distinct mechanisms that one aggregate Elo number cannot separate. 7 essays · 4 frozen · 3 design.
A chess move arrives as one action, but producing it requires several different computations.
Two agents lose to the same opponent.
A model can be cheap on a price sheet and still be impractical inside an agent.
The easiest response to a weak chess-playing language model is to give it more chess text, more tools, more reflection, and more tokens.
Two chess agents can have the same unaided score and still be very different reasoning systems.
Suppose two agents find Stockfish's preferred move 31.5% of the time.
Here is a result that practically writes its own misleading headline.
Path 03 / Design
Follow the benchmark, pre-label audit, and preregistered decisions that define what a later outcome would mean. 5 essays · 5 design.
Release status: v0.1 draft. This document specifies a benchmark; it does not announce a terminal benchmark release. The Luna and DeepSeek ladders are exploratory pilot evidence, not v1 leaderboard submissions.
The dangerous moment in an experiment is not always when the result appears. Sometimes it is the moment just before an expensive job starts, when the code looks complete, the sample is frozen, and everyone is t
The first games looked absurdly easy.
The most misleading condition in our next experiment is called no_information.
We had already designed the intervention.
Path 04 / Audit
Follow the byte-level, transport, label, and no-result gates that stop a working protocol from becoming a stronger claim than its evidence. 5 essays · 5 design.
A105 began with an apparently simple promise: change whether a model can read a board panel or a complete legal-action panel, and change nothing else.
The previous systems audit ended at an intentionally artificial boundary. A105-R2 proved that our repository adapter could carry 128 canonical treatment cells into two in-memory serialization profiles without a
A failure discovered before outcomes creates a temptation: fix the code, rerun the test, and describe the system as if the failure never happened.
It is easy to describe position labeling as housekeeping. Collect some chess positions, ask a strong engine whether each one is tactical, ambiguous, or quiet, and balance the final benchmark across those catego
There is a satisfying kind of research progress that produces no number.
Path 05 / Chronology
Track operations, shuffled inputs, transport repairs, and later results in numbered order; the evidence badge remains the public boundary. 9 essays · 3 frozen · 6 provisional.
At one point in our DeepSeek chess run, two numbers seemed impossible to reconcile. The token counter was reported as 19.277 million. The runtime was described as roughly 117 hours. Yet the experiment had not b
At the 2026-08-04 drafting boundary, the most important output of our DeepSeek chess evaluation did not exist.
The most tempting graph in this project slopes the wrong way.
The naive experiment is easy: show one list, shuffle it, ask twice, and count changes. It is also wrong. A model can change its answer even when the prompt is byte-for-byte identical. Without measuring that bas
We tried to measure a prompt-guidance effect twice, paid for sixteen chat intents, and refused to read either attempt as an experiment.
E026's live transport aborted before any chat return; E027's chat transport succeeded for ten returns, but metadata attestation failed on the immediate lookup. E028 then ran as a fresh response cohort with zero
An exact answer from a single frozen source does not make an experiment exact.
E029 asked a narrower question than a game ladder can answer. If Luna receives the same board but also sees two unranked moves drawn from an exact Syzygy policy, does its submitted move preserve more 50-move-ru
The sharp question was not whether an engine can play chess. It was whether a small engine-generated action scaffold changes what the same language model can do when everything else stays fixed.