The Curve That Runs Backward
Two model-plus-harness systems currently look stronger when shown fewer engine candidates. That is a replicated descriptive shape—not yet a causal effect.
The most tempting graph in this project slopes the wrong way.
If a chess agent receives ten plausible moves, it has more opportunities than if it receives only two. The best two moves have not vanished from the larger list. A naive “more help cannot hurt” account therefore predicts flat or improving play as the list widens.
Our observed point estimates reverse that prediction. In Luna's frozen pilot, the candidate-assisted conditions order top-2, top-3, top-5, then top-10. The unfinished DeepSeek V4 Flash replication currently has the same ordering.
That agreement is interesting. It is not the causal conclusion.
This distinction matters because a ladder is an excellent instrument for finding a model's approximate operating range and a poor instrument for isolating why one interface beats another. The current curves combine candidate realization, comparison load, different games, adaptive opponents, and long trajectory effects. A second system showing the same descriptive shape raises the priority of the mechanism experiment. It does not make the missing controls appear retroactively.
The observation, with its evidence states attached
The current synthesis joins two unlike evidence objects. Luna is terminal and content-addressed. DeepSeek is an interim repeated look at a cohort whose ladders are still running.
| Cohort | State | Variant | k | Games | W-D-L | Protocol rating | Interval |
|---|---|---|---|---|---|---|---|
| Luna | terminal, frozen | top-10 | 10 | 10 | 2-2-6 | 1126.7 | 804.5-1407.3 |
| Luna | terminal, frozen | top-5 | 5 | 60 | 28-14-18 | 1830.1 | 1712.8-1949.8 |
| Luna | terminal, frozen | top-3 | 3 | 80 | 43-22-15 | 2430.9 | 2317.4-2547.7 |
| Luna | terminal, frozen | top-2 | 2 | 60 | 41-11-8 | 2832.4 | 2680.0-2991.8 |
| DeepSeek | interim, unfrozen | top-10 | 10 | 10 | 3-1-6 | 1247.5 | 998.2-1466.2 |
| DeepSeek | interim, unfrozen | top-5 | 5 | 10 | 5-1-4 | 1389.9 | 1171.0-1618.0 |
| DeepSeek | interim, unfrozen | top-3 | 3 | 12 | 7-1-4 | 1479.2 | 1276.1-1703.0 |
| DeepSeek | interim, unfrozen | top-2 | 2 | 14 | 12-0-2 | 1851.6 | 1596.2-2190.0 |
Every displayed number belongs to the complete system: model, candidate generator, prompt, parser, tool contract, reasoning setting, color, Stockfish 18 opponent, and adaptive ladder rule. These are local, Black-only protocol ratings—not universal base-model Elo.
The Luna Davidson-WDL estimates use all frozen anchors. Their profile intervals are ordinary marginal intervals; they are neither optional-stopping-corrected nor multiplicity-adjusted pairwise tests. The DeepSeek values are more fragile: they are fractional-score logistic quasi-Elo summaries from an unfinished adaptive cohort. Every adjacent DeepSeek working interval overlaps. The clean ordering of its point estimates looks much more certain than the evidence is.
Even the no-candidate row does not extend the curve to k=0. Regular play uses
a different interface in which the model must propose a move rather than select
from an oracle-generated consideration set. It is a useful baseline for the
whole system, but not another width on the same selection treatment.
What replicated—and what did not
One fact currently replicates at the descriptive level:
Among candidate-assisted variants, both cohorts' observed point estimates rise at every step as candidate width falls from 10 to 5 to 3 to 2.
That sentence deliberately has four limiters: candidate-assisted, observed, point estimates, and currently.
It does not say that two candidates cause better decisions than ten. It does not establish a terminal DeepSeek leaderboard. It does not establish that the same mechanism produced both shapes. And it does not license a numerical Luna-versus-DeepSeek ranking, because one side is frozen while the other is still moving.
The repeated shape does change our research priorities. Sampling noise remains possible, especially in small interim blocks, but “the first curve was merely a one-model curiosity” is now less satisfying as the only explanation. The right response is not stronger prose. It is a design in which the width comparison is actually paired.
Why the ladders cannot identify candidate-count causality
The candidate sets were not nested
The pilot conditions were generated separately. The top-2 list in one game was
not guaranteed to be the first two members of the top-10 list in a matched copy
of the same position. Changing k could therefore change both the burden of
comparison and the realized quality of the available actions.
If top-2 happened to receive cleaner engine proposals, its advantage would not measure overload. Conversely, a top-10 list can contain the same strong core plus eight plausible distractors; failure there could reflect comparison load, rank or order sensitivity, extra tokens, or a changed reasoning trajectory. The present data cannot separate these stories.
The games were not paired
Each condition played its own games from the initial position. A move at ply 40 is conditioned on 39 previous moves made by that treatment. Different candidate widths therefore encounter different positions, tactical hazards, history lengths, and recovery opportunities.
Full games tell us whether a system survives the consequences of its own actions. That is exactly why they are valuable. It is also why their outcomes cannot isolate a local selection effect without matched positions.
The opponent anchors were adaptive
Strong early blocks promoted a variant to harder opponents; middling blocks triggered repeats; weak blocks stopped. This is efficient bracketing, but it creates unequal sample sizes and data-dependent anchor paths. The frozen Luna rating model uses all completed anchors, yet its marginal intervals do not undo the exploratory path. DeepSeek adds repeated looks while suites are live.
The ladder answers “where should we test next?” It was not designed to answer “what is the causal effect of removing three candidates while everything else is fixed?”
The protocol holds color and opening fixed
The LLM always played Black from the standard initial position. Candidate order was randomized and scores and ranks were hidden, but one color and one opening family remain a narrow slice of chess. The inverse curve may be specific to this prompt, this candidate generator, this time control, or this provider snapshot. A protocol-specific regularity is still worth studying. It should not be renamed a model trait before it travels.
The experiment that can turn shape into effect
The decisive next step is a nested-set paired candidate-count crossover.
For each frozen test position, one engine pass constructs a canonical top-10
parent set. The k=2, k=3, and k=5 conditions are strict nested subsets of
that same parent. Every condition sees the same root position in a fresh
context, under the same reasoning budget and action schema.
The design needs more than shared FENs:
- Pair every width within position. Run
k in {2, 3, 5, 10}on every eligible position rather than assigning different positions to widths. - Repeat presentation orders. Shuffle each displayed set with registered seeds, never expose engine ranks or scores, and balance candidate slots so a width effect cannot be a primacy effect in disguise.
- Use both colors and isolated source families. Balance side to move and cluster uncertainty by source game; keep normalized and color-mirrored families in one partition.
- Score the decision, not only survival. Record legality, chosen source rank, independent-evaluator WDL regret, catastrophic-blunder rate, response stability under order permutations, tokens, latency, and protocol failure.
- Freeze stopping and contrasts before outcomes. The primary adjacent
contrasts are
2-3,3-5, and5-10, with source-game clustered uncertainty and a declared multiplicity rule. No adaptive anchor is needed for a position-level comparison.
Nested sets isolate a meaningful intervention: what happens when additional known candidates are appended to the same strong core. They do not, by themselves, explain why performance changes. A second decomposition must vary candidate provenance and distractor quality at fixed width: model-proposed moves, oracle-included sets, best-move-omitted sets, controlled legal distractors, and all-legal sets. That separates proposal coverage from selection load and tests whether “fewer is better” survives when quality is controlled.
The full-game ladder and paired position study then play complementary roles. The ladder measures end-to-end survival. The paired crossover measures the local consequence of widening one consideration set. Agreement between them would be far stronger than either alone; disagreement would reveal that trajectory dynamics, not local choice overload, produced the Elo curve.
What would change our mind
The comparison-overload interpretation should weaken if any of the following occurs:
- the terminal DeepSeek freeze changes or erases the interim ordering;
- paired nested sets show no adverse
k=10contrast after candidate identity, position, order, color, and budget are controlled; - the apparent width effect vanishes when candidate quality is held fixed;
- order permutations explain most of the regret difference;
- the effect appears only in full-history games and not in fresh-position decisions; or
- prospective fixed-anchor, balanced-color games fail to reproduce the Luna ordering.
The opposite result would also be informative. If the same model chooses well from a nested top-2 core but degrades as matched distractors are added—and that effect survives order counterbalancing, both colors, independent evaluation, and source-game clustering—then “comparison load” becomes an empirical mechanism rather than a name for a surprising graph.
The research claim today
Today we have one terminal curve and one interim echo.
The terminal claim is narrow: within Luna's frozen, Black-only candidate-assisted
pilot, every joint-rating point estimate increases as k falls from 10 to 2.
The interim claim is narrower still: DeepSeek's current unfrozen point estimates
show the same ordering, with overlapping adjacent working intervals.
The scientific opportunity is larger than either claim. Candidate width may be a controlled probe of how an agent converts proposals into action. But the fingerprint is not the backward line already on the chart. The fingerprint is the response that remains after positions, candidate identity, order, color, budget, stopping, and evaluation have stopped moving underneath it.
Evidence and reproduction
This essay is bound to the read-only ladder synthesis with analysis identity
2ff055248abedab01a6f1b4eeddb657f4f9b2098b6a2fa1ca67cc7a7f3cd0baa.
The source JSON file SHA-256 is
67563cc19bedf1eee9dada44289173383d51854bd224a608e13f1a8155f16994.
DeepSeek values are the interim values in that artifact, not a live refresh and
not a terminal freeze.
uv run python research/analysis/ladder_synthesis.py
uv run pytest -q tests/test_ladder_synthesis.py \
tests/test_blog23_inverse_candidate_curve_story.py
- Full-game ladder synthesis
- Machine-readable synthesis
- E020/E021 preregistration
- Candidate curves as model fingerprints
Operational paths, raw logs, credentials, mutable run metadata, and unpublished artifacts are intentionally excluded. Research references resolve to the versioned publication bundle rather than the host filesystem.