Evidence / full-game protocol
Protocol leaderboard
Every result has its protocol attached: the model plays Black only in full games against Stockfish 18 with UCI LimitStrength. Read each cohort's evidence state separately; these local measurements are not a cross-cohort ranking.
Scope. Black-only model games · Stockfish 18 with UCI LimitStrength · 200-ply limit. W-D-L and scores include clean completed games only; provider-error records are excluded.
Ranked by joint local Elo
Luna
Top 3
3 randomized, unranked proposalsTop 5
5 randomized, unranked proposalsTop 10
10 randomized, unranked proposalsRegular
No engine proposalsNon-comparable cohorts. Do not rank the cohort values against one another. These are local protocol-specific measurements, not FIDE, online-platform, or general model ratings.
Candidate-width fingerprint
The curve runs backward.
Among candidate-assisted conditions, the observed point estimate rises as the visible set narrows. Select a point to inspect the uncertainty—not just the line.
Table / text view
This native table lists every cohort and variant represented by the curve data, including the separate regular interface. Frozen cohorts report local protocol Elo and profile 95% intervals; interim cohorts report working estimates and working 95% intervals. “Not published” identifies a missing estimate; “No interval or bound published” identifies missing uncertainty data.
| Model / cohort | Variant | Candidate count | Estimate value | 95% interval or bound | Estimate games | Estimate W-D-L | Evidence state |
|---|---|---|---|---|---|---|---|
| Luna | Regular interface (regular) | Separate interface | 1067 | 705–1351 | 10 | 0–5–5 | Frozen |
| Luna | Top 10 (top10) | 10 | 1127 | 805–1407 | 10 | 2–2–6 | Frozen |
| Luna | Top 5 (top5) | 5 | 1830 | 1713–1950 | 60 | 28–14–18 | Frozen |
| Luna | Top 3 (top3) | 3 | 2431 | 2317–2548 | 80 | 43–22–15 | Frozen |
| Luna | Top 2 (top2) | 2 | 2832 | 2680–2992 | 60 | 41–11–8 | Frozen |
| DeepSeek V4 Flash | Regular interface (regular) | Separate interface | 973 | 467–1264 | 10 | 0–2–8 | Interim working |
| DeepSeek V4 Flash | Top 10 (top10) | 10 | 1248 | 998–1466 | 10 | 3–1–6 | Interim working |
| DeepSeek V4 Flash | Top 5 (top5) | 5 | 1390 | 1171–1618 | 10 | 5–1–4 | Interim working |
| DeepSeek V4 Flash | Top 3 (top3) | 3 | 1479 | 1276–1703 | 12 | 7–1–4 | Interim working |
| DeepSeek V4 Flash | Top 2 (top2) | 2 | 1852 | 1596–2190 | 14 | 12–0–2 | Interim working |
Frozen Luna evidence. These protocol-specific estimates do not identify a candidate-count effect: sets were separately generated, games were unpaired, and anchors were adaptive.
Protocol scope
Black-only
full games.
- Opponent
- Stockfish 18 with UCI LimitStrength
- Model color
- Black only; no color-balanced claim
- Clock
- 0.1 seconds per Stockfish move
- Game limit
- 200 plies maximum
- Reasoning
- high effort
- Candidate display
- randomized; scores and ranks hidden
Interpretation limits
Read each number with its boundary.
local scale — These are local protocol-specific measurements, not FIDE, online-platform, or general model ratings.
adaptive design — Anchors were selected adaptively. Earlier blocks remain part of the frozen Luna likelihood.
interim boundary — DeepSeek terminal data and publication authorization are separate states: terminal rows can be shown while the freeze-bound rating remains withheld until current release gates authorize publication.
causal boundary — Candidate conditions are complete agent protocols, not isolated estimates of base-model chess strength.