← All writing
Essay 09 / 22Frozen evidence

Candidate Curves as Model Fingerprints

A model may be defined as much by the help it can use as by what it can do alone

Two chess agents can have the same unaided score and still be very different reasoning systems.

One may improve when given ten plausible moves. Another may need the list compressed to two. One may reliably compare engine proposals but fail to generate them. Another may generate the right move and then talk itself out of playing it. Their unaided Elo compresses those differences into one number.

A candidate curve asks a richer question:

How does the complete system respond as proposal bandwidth, proposal quality, presentation, and inference budget change?

Our completed Luna pilot supplies one striking curve. A live DeepSeek replication already looks different in raw outcomes and inference volume. It does not yet supply a second terminal curve. This post therefore has two jobs: preserve the dated observation we actually have, and specify the cross-model experiment that could turn candidate response into a defensible model fingerprint.

The thesis is a hypothesis, not a result:

A vector of controlled assistance responses may distinguish model families more reliably than unaided full-game Elo alone.

Evidence state at a glance

Three evidence states must not be mixed.

Completed and frozen

The Luna pilot contains 220 clean games and has a content-addressed freeze. Its joint all-anchor estimates are reproducible from that fixed cohort.

Luna conditionClean gamesW-D-LJoint local EloProfile 95%
No candidates100–5–51067705–1351
Top 10102–2–61127805–1407
Top 56028–14–1818301713–1950
Top 38043–22–1524312317–2548
Top 26041–11–828322680–2992

These are Black-only, Stockfish-18-anchored ratings under this protocol. The candidate sets were produced by separate time-limited searches, the anchors were adaptive, and the games were unpaired. Luna's shape is a completed exploratory systems curve, not a causal measurement of candidate count.

Live and provisional

At the DeepSeek efficiency snapshot dated 2026-08-04T21:10:33Z, the canonical ledger contained 43 clean games and zero provider-error games:

DeepSeek conditionClean gamesRaw W-D-LRecorded tokens / clean gameState
No candidates100–2–8759,7231320 block complete
Top 1083–0–5610,8801320 block active
Top 563–1–2924,3491320 block active
Top 386–0–2368,8291320 block active
Top 21110–0–1313,6511320 complete; 1720 active

The top-2 total combines a 10–0–0 score at 1320 with one completed loss at 1720. The 10–0 block has no finite maximum-likelihood Elo. It is evidence that the anchor was too weak and that the ladder should promote—not permission to print a saturated point estimate.

Nor are the raw scores a candidate curve. Conditions have unequal sample sizes, top-2 has reached another anchor, four suites were active at the dated snapshot, and unfinished games have already consumed time without entering the table. No terminal DeepSeek rating curve or cross-model ranking is claimed here. The zero-error statement is limited to persisted terminal game records at this dated cut. It does not observe request-level retries, requests without game JSON, or route-specific failures; A103 reports that missingness boundary separately.

Planned and gated

The fingerprint claim requires nested paired candidates, both colors, frozen openings, matched inference budgets, repeated presentation orders, terminal cohorts, and at least two additional independent model families. Those data do not exist yet.

What the Luna shape suggests

Luna's realized curve contains more structure than “assistance helped.”

  • Top-10 remained near unaided play within wide uncertainty.
  • Top-5 produced a large strength increase.
  • Top-3 improved again.
  • Top-2 was strongest and used the fewest aggregate tokens per clean game.

This shape is consistent with a system that benefits from oracle proposal coverage but pays a steep comparison or search cost as the consideration set widens. It is also consistent with non-nested candidate realization, correlated game trajectories, adaptive anchors, color effects, or sampling noise.

The fingerprint proposal starts only after those alternatives are controlled. A fingerprint is not the five pilot Elo values. It is a reproducible response profile under a preregistered intervention family.

A fingerprint is a vector, not an optimum

Calling a model “best at top-2” throws away most of the useful information. The registered fingerprint should contain features from four layers.

1. Strength-response shape

For candidate widths k ∈ {0, 2, 3, 5, 10}, estimate on held-out paired positions and replicated full games:

  • ΔR(k): rating or paired-regret improvement relative to unaided play;
  • k*: width with the best posterior expected performance;
  • overload_2_10 = R(2) − R(10);
  • local slopes between adjacent widths;
  • discrete curvature and the number of monotonicity reversals;
  • area under the normalized assistance-response curve; and
  • uncertainty that k* is truly optimal rather than a sampling winner.

The raw rating level remains a separate feature. Otherwise a weak model helped to mediocrity could look identical to a strong model helped to mastery.

2. Proposal and selection decomposition

Paired diagnostics should measure:

  • unaided recall of the engine-best move;
  • best-move coverage of the model's self-generated shortlist;
  • regret when selecting from an oracle shortlist;
  • chosen source rank and value loss from the best available candidate;
  • performance when the best candidate is deliberately omitted;
  • engine-selects-model-proposals versus model-selects-engine-proposals; and
  • value-ranking calibration on successor positions.

These features separate a missing policy from a weak critic. Two models with the same top-2 Elo lift may arrive there for opposite reasons.

3. Stability and failure geometry

The same candidate set should be shown under repeated deterministic order permutations. Record:

  • probability that the chosen move changes under order permutation;
  • variance of regret across order seeds;
  • illegal-action and tool-protocol failure rates;
  • catastrophic-blunder hazard by game phase;
  • sensitivity to full history versus a fresh position; and
  • failure-label distribution across state, tactic, value, plan, memory, and interface errors.

A model that is strong on average but highly order-sensitive has a different agent profile from one that is slightly weaker and stable.

4. Compute-response shape

For matched low, medium, and high reasoning budgets, record:

  • tokens and latency per decision;
  • tokens, wall time, and worker-hours per completed game;
  • improvement per additional thousand completion-side tokens;
  • candidate-width × reasoning-budget interaction;
  • retry and correction turns per legal move; and
  • strength-efficiency Pareto membership.

This layer matters because two systems can reach the same move quality with very different sequential depth and operating cost. The dated DeepSeek snapshot is already an operational warning: 43 games consumed 24.43 million recorded tokens and 157.55 worker-hours. It is not yet a terminal efficiency fingerprint.

Normalize without erasing the phenomenon

Raw Elo differences cannot be compared naively across models. A stronger unaided model has less headroom, and performance near the top or bottom of an anchor range saturates. Several representations should therefore be published together:

  1. absolute protocol-specific strength with W-D-L and uncertainty;
  2. within-model lift relative to the unaided condition;
  3. paired position regret on a shared continuous engine-value scale;
  4. catastrophic-blunder probability;
  5. compute per decision at matched positions; and
  6. rank-based curve shape, which is less sensitive to rating scale.

No single normalization is privileged. If the apparent family separation exists only after one convenient transformation, it is not a robust fingerprint.

Measuring distance between fingerprints

Let f_m be the preregistered feature vector for model m. A plain Euclidean distance would be misleading: Elo, token counts, order sensitivity, and regret live on different scales and have different uncertainties.

The primary distance should be an uncertainty-aware standardized distance:

d²(i, j) = (f_i − f_j)ᵀ [Σ_population + Σ_i + Σ_j + λI]⁻¹ (f_i − f_j)

Σ_i and Σ_j are bootstrap or posterior covariance matrices for the two fingerprints. Σ_population scales features by variation observed across the registered model panel. λI is a preregistered shrinkage term that prevents a small pilot panel from producing an unstable inverse.

That omnibus distance should be accompanied by interpretable component distances:

  • inverse-variance-weighted integrated distance between normalized k curves;
  • Jensen-Shannon distance between failure-label distributions;
  • absolute difference in order-instability probability;
  • log-ratio distance for tokens and latency; and
  • difference in proposal-versus-selection regret.

Bootstrap the entire experiment by opening/position block, not by treating every move as independent. Report a confidence interval for each pairwise distance and the within-model test-retest distance. A between-model difference is meaningful only when it reliably exceeds the same model's variation across seeds, days, openings, and provider snapshots.

Clustering is exploratory. The confirmatory question is not “does a dendrogram look interesting?” It is whether a fingerprint measured on one position split predicts held-out behavior: the best candidate width, order sensitivity, proposal-selection asymmetry, and benefit from uncertainty-triggered search.

The matched cross-model protocol

Every model family must face the same frozen experimental object.

Candidate construction

  • Run one Stockfish 18 fixed-node max-10 analysis per position.
  • Cache the exact engine binary checksum, options, nodes, candidates, scores, and principal variations.
  • Derive top-2, top-3, top-5, and top-10 as literal prefixes.
  • Hide ranks, scores, and principal variations in the move-only condition.
  • Apply the same candidate-order seeds to every model.

Positions and games

  • Use a frozen opening suite and paired diagnostic positions unseen in the exploratory games.
  • Balance opening, middlegame, and endgame; tactical, ambiguous, and quiet; White and Black to move.
  • Run color-reversed mini-matches from identical openings.
  • Use independent engine adjudication rather than treating every 200-ply cutoff as an unquestioned draw.
  • Allocate samples by a preregistered rule and fit all anchors jointly.

Model interface

  • Pin prompt bytes, tool schemas, legal-action encoding, retry limits, context policy, and reasoning budget.
  • Record requested and returned model/provider identifiers on every request.
  • Interleave families and conditions so provider day and machine load are not confounded with treatment.
  • Match maximum output and reasoning budgets; publish when an API cannot express an equivalent control.
  • Preserve request-level timestamps, token classes, retries, and billed activity join keys.

The same semantic task is not enough. A different adapter, hidden system prompt, or constrained-output mechanism can change the player.

Required model panel

Luna and DeepSeek are two observations, not a taxonomy. The minimum publishable panel has four independent families:

  1. the completed Luna family;
  2. the terminal DeepSeek family;
  3. a frontier model from a third provider and training lineage; and
  4. an open-weight family with a pinned checkpoint and reproducible inference stack.

At least one family should include matched reasoning and non-reasoning siblings, or the same checkpoint under genuinely enforceable reasoning budgets. At least one should be rerun through two interface implementations to estimate adapter variance. Each family needs two temporally separated replications or immutable snapshot identifiers.

This is the minimum needed to distinguish “model family” from “two APIs behaved differently during one week.” A stronger benchmark would add a compact model, a second open-weight lineage, and a chess-specialized policy/value baseline.

Analysis gates

The article advances in stages.

Gate A: terminal descriptive replication

  • Every variant reaches its registered stopping condition.
  • Provider-error games are excluded and separately reported.
  • Raw W-D-L, anchors, tokens, and timing reconcile to a frozen cohort manifest.
  • Saturated blocks receive one-sided bounds, never invented point estimates.

Passing Gate A licenses a completed DeepSeek systems curve. It does not license a mechanistic fingerprint claim.

Gate B: paired causal candidate curves

  • Candidate sets are cached and nested.
  • Every family completes the same paired position cells.
  • Order, color, phase, and position-stratum effects are estimated.
  • Candidate value gaps and best-move coverage are controlled.

Passing Gate B licenses statements about response to proposal bandwidth under the registered interface.

Gate C: test-retest reliability

  • Within-model fingerprints replicate across seeds, days, and held-out openings.
  • Between-family distances exceed within-family distances with uncertainty.
  • The feature set predicts held-out assistance response better than unaided Elo alone.

Only Gate C licenses “candidate curves act as model fingerprints.”

Gate D: architectural prediction

  • Fingerprints predict which models benefit from shortlist compression, branch-isolated comparison, or uncertainty-triggered widening.
  • Interventions selected from the first split improve held-out games without merely increasing oracle leakage.

Gate D turns a descriptive fingerprint into an engineering instrument.

Falsifiers

The fingerprint hypothesis becomes less credible if any of the following hold:

  • nested paired candidates eliminate the Luna top-10 disadvantage;
  • within-model distances across order seeds or provider dates are as large as between-family distances;
  • curves cluster by adapter, prompt length, or output constraint rather than by model family;
  • the apparent separation disappears after controlling unaided strength and candidate best-move coverage;
  • reasoning-budget changes move a model farther than changing model family;
  • family labels cannot be predicted on held-out positions above a preregistered baseline;
  • fingerprint features do not predict the best held-out scaffold; or
  • adding independent families destroys the original clustering.

These are not nuisances to explain away. They distinguish a model property from an experimental artifact.

What would change our mind?

The strongest version of the idea—that each family has a stable characteristic curve—fails if the curve is mostly an interface fingerprint. In that case the useful unit of analysis is model-plus-adapter, which would still matter for agent design but would not support claims about latent model architecture.

The Luna interpretation changes if its paired confirmatory curve becomes flat or monotonic. The DeepSeek interpretation changes with every completed live game until its cohort is terminal and frozen. A different optimal k in DeepSeek would be interesting; the same optimal k would not establish a shared mechanism.

The broader project succeeds even if candidate curves do not classify model families. A reliable curve could still tell us how to allocate proposal breadth, comparison depth, and compute for one deployed agent. Prediction, not a pretty cluster plot, is the standard.

Evidence required before publication

The terminal version of this post must link:

  • the Luna and DeepSeek content-addressed cohort freezes;
  • freezes for at least two additional model families;
  • exact prompts, tool schemas, returned model IDs, and provider metadata;
  • the cached nested candidate manifest and engine checksum;
  • preregistered sample allocation, exclusions, and distance metric;
  • per-condition raw W-D-L, anchors, and uncertainty;
  • paired regret, coverage, order-stability, and failure endpoints;
  • request-level tokens, latency, retries, and reconciled billing;
  • bootstrap code and fingerprint covariance matrices;
  • held-out prediction results and negative controls; and
  • the publication-bundle manifest with no missing, stale, live, or secret-bearing evidence.

Until then, the completed Luna curve is an observation, the DeepSeek table is a dated live snapshot, and the model-fingerprint thesis is a registered bet.

Public-surface note

Operational paths, raw logs, credentials, mutable run metadata, and unpublished artifacts are intentionally excluded. Research references resolve to the versioned publication bundle rather than the host filesystem.