Public notebook · 30 essays · evidence snapshot

Read the chain.

Choose an entry point that matches your question, or follow the numbered sequence. Every essay remains available once, with its current public evidence state attached.

Public evidence snapshotAugust 6, 2026 at 6:02 AM UTC

Snapshot chronology / public boundary

Order tells the story.
State tells the boundary.

01

Essay numbers preserve research order, not publication dates. The state badge is the public evidence boundary for the essay; it does not upgrade a design or interim note because a later result exists.

Three clocks for one evaluation program. Snapshot intervals provide observation bounds, not exact process launch or exit timestamps. No account usage, credential balance, or model-specific dollar attribution is included. These clocks describe evaluation operations and do not support a strength, causal, or terminal DeepSeek claim.

Cutoff August 6, 2026 at 6:02 AM UTC · source identity d6b9d1726427

Evidence key / 30 essays

Three states.
One vocabulary.

02
01
Frozen evidence

11 frozen essays

A bounded observation tied to frozen public evidence; it is not automatically causal or general.

02
Provisional

6 provisional essays

A dated or interim view; it is not a terminal result.

03
Protocol / design

13 design essays

A protocol, audit, or prospective method; it carries no outcome by itself.

Reading paths / no duplicate links

Start where the
question is.

03

These paths are thematic entry points through the same numbered chronology. Each public essay appears in one path only; every link keeps its canonical slug.

01Measurement4 essays02Diagnosis7 essays03Design5 essays04Audit5 essays05Chronology9 essays

Path 01 / Measurement

Start with the measurement

04

Begin with the surprising result, then attach the full-game protocol before reading a rating as a property of a model. 4 essays · 4 frozen.

Path 02 / Diagnosis

Find the failure inside the score

05

Trace candidate assistance, interface behavior, and the distinct mechanisms that one aggregate Elo number cannot separate. 7 essays · 4 frozen · 3 design.

Path 03 / Design

Make the intervention falsifiable

06

Follow the benchmark, pre-label audit, and preregistered decisions that define what a later outcome would mean. 5 essays · 5 design.

Path 04 / Audit

Protect the evidence path

07

Follow the byte-level, transport, label, and no-result gates that stop a working protocol from becoming a stronger claim than its evidence. 5 essays · 5 design.

Path 05 / Chronology

Read the later snapshot without upgrading it

08

Track operations, shuffled inputs, transport repairs, and later results in numbered order; the evidence badge remains the public boundary. 9 essays · 3 frozen · 6 provisional.

14
Laboratory operations

Five Workers, One Clock: Why LLM Chess Runs Feel Slow

At one point in our DeepSeek chess run, two numbers seemed impossible to reconcile. The token counter was reported as 19.277 million. The runtime was described as roughly 117 hours. Yet the experiment had not b

Provisional
9 min read1,951 words
15
Laboratory operations

The Result That Refused to Exist

At the 2026-08-04 drafting boundary, the most important output of our DeepSeek chess evaluation did not exist.

Provisional
11 min read2,404 words
23
Laboratory operations

The Curve That Runs Backward

The most tempting graph in this project slopes the wrong way.

Provisional
8 min read1,659 words
24
Laboratory operations

The Shuffle That Looked Like a Signal

The naive experiment is easy: show one list, shuffle it, ask twice, and count changes. It is also wrong. A model can change its answer even when the prompt is byte-for-byte identical. Without measuring that bas

Frozen evidence
2 min read446 words
26
Laboratory operations

The Result That Arrived Late

We tried to measure a prompt-guidance effect twice, paid for sixteen chat intents, and refused to read either attempt as an experiment.

Frozen evidence
4 min read833 words
27
Laboratory operations

The Prompt Did Not Whisper

E026's live transport aborted before any chat return; E027's chat transport succeeded for ten returns, but metadata attestation failed on the immediate lookup. E028 then ran as a fresh response cohort with zero

Frozen evidence
5 min read955 words
28
Laboratory operations

The Single-Source Oracle Was Exact. The Protocol Wasn't.

An exact answer from a single frozen source does not make an experiment exact.

Provisional
5 min read1,128 words
29
Laboratory operations

Two Exact Moves Changed the Submitted Policy

E029 asked a narrower question than a game ladder can answer. If Luna receives the same board but also sees two unranked moves drawn from an exact Syzygy policy, does its submitted move preserve more 50-move-ru

Provisional
4 min read695 words
30
Laboratory operations

Two Unranked Engine Moves Helped Luna Survive

The sharp question was not whether an engine can play chess. It was whether a small engine-generated action scaffold changes what the same language model can do when everything else stays fixed.

Provisional
3 min read673 words