← All writing
Essay 23 / 30Provisional

The Curve That Runs Backward

The 2026-08-05 synthesis snapshot shows a descriptive ordering in which two model-plus-harness systems look stronger when shown fewer engine candidates. It is not a causal effect.

The most tempting graph in this project slopes the wrong way.

If a chess agent receives ten plausible moves, it has more opportunities than if it receives only two. The best two moves have not vanished from the larger list. A naive “more help cannot hurt” account therefore predicts flat or improving play as the list widens.

In the 2026-08-05 synthesis snapshot, Luna's frozen pilot ordered the candidate-assisted conditions top-2, top-3, top-5, then top-10. The then- unfinished DeepSeek V4 Flash replication had the same ordering in its interim point estimates.

That agreement is interesting. It is not the causal conclusion.

This distinction matters because a ladder is an excellent instrument for finding a model's approximate operating range and a poor instrument for isolating why one interface beats another. The snapshot's curves combine candidate realization, comparison load, different games, adaptive opponents, and long trajectory effects. A second system showing the same descriptive shape raised the priority of the mechanism experiment. It did not make the missing controls appear retroactively.

The 2026-08-05 observation, with its evidence states attached

The 2026-08-05 synthesis joined two unlike evidence objects. Luna was terminal and content-addressed. DeepSeek was an interim repeated look at a cohort whose ladders were still running at the cutoff.

If columns extend beyond the viewport, focus this region and use the left and right arrow keys to scroll horizontally.
CohortStateVariantkGamesW-D-LProtocol ratingInterval
Lunaterminal, frozentop-1010102-2-61126.7804.5-1407.3
Lunaterminal, frozentop-556028-14-181830.11712.8-1949.8
Lunaterminal, frozentop-338043-22-152430.92317.4-2547.7
Lunaterminal, frozentop-226041-11-82832.42680.0-2991.8
DeepSeekinterim, unfrozentop-1010103-1-61247.5998.2-1466.2
DeepSeekinterim, unfrozentop-55105-1-41389.91171.0-1618.0
DeepSeekinterim, unfrozentop-33127-1-41479.21276.1-1703.0
DeepSeekinterim, unfrozentop-221412-0-21851.61596.2-2190.0

Every displayed number in this snapshot belongs to the complete system: model, candidate generator, prompt, parser, tool contract, reasoning setting, color, Stockfish 18 opponent, and adaptive ladder rule. These are local, Black-only protocol ratings—not universal base-model Elo.

The Luna Davidson-WDL estimates use all frozen anchors. Their profile intervals are ordinary marginal intervals; they are neither optional-stopping-corrected nor multiplicity-adjusted pairwise tests. The DeepSeek values are more fragile: they were fractional-score logistic quasi-Elo summaries from an unfinished adaptive cohort at the cutoff. Every adjacent DeepSeek working interval overlapped. The clean ordering of its point estimates looked much more certain than the evidence was.

Even the no-candidate row does not extend the curve to k=0. Regular play uses a different interface in which the model must propose a move rather than select from an oracle-generated consideration set. It is a useful baseline for the whole system, but not another width on the same selection treatment.

What the snapshot shows—and what it does not

The same descriptive ordering appears in this snapshot: among candidate-assisted variants, both cohorts' observed point estimates rose at every step as candidate width fell from 10 to 5 to 3 to 2. That does not say that two candidates caused better decisions than ten, establish a terminal DeepSeek leaderboard, identify a shared mechanism, or license a numerical Luna-versus-DeepSeek ranking; Luna was frozen while DeepSeek was interim at the 2026-08-05 cutoff. It raised the priority of paired mechanism work, not the strength of the claim.

Why the ladders cannot identify candidate-count causality

The candidate sets were not nested

The pilot conditions were generated separately. The top-2 list in one game was not guaranteed to be the first two members of the top-10 list in a matched copy of the same position. Changing k could therefore change both the burden of comparison and the realized quality of the available actions.

If top-2 happened to receive cleaner engine proposals, its advantage would not measure overload. Conversely, a top-10 list can contain the same strong core plus eight plausible distractors; failure there could reflect comparison load, rank or order sensitivity, extra tokens, or a changed reasoning trajectory. The present data cannot separate these stories.

The games were not paired

Each condition played its own games from the initial position. A move at ply 40 is conditioned on 39 previous moves made by that treatment. Different candidate widths therefore encounter different positions, tactical hazards, history lengths, and recovery opportunities.

Full games tell us whether a system survives the consequences of its own actions. That is exactly why they are valuable. It is also why their outcomes cannot isolate a local selection effect without matched positions.

The opponent anchors were adaptive

Strong early blocks promoted a variant to harder opponents; middling blocks triggered repeats; weak blocks stopped. This is efficient bracketing, but it creates unequal sample sizes and data-dependent anchor paths. The frozen Luna rating model uses all completed anchors, yet its marginal intervals do not undo the exploratory path. DeepSeek added repeated looks while suites were live at the snapshot cutoff.

The ladder answers “where should we test next?” It was not designed to answer “what is the causal effect of removing three candidates while everything else is fixed?”

The protocol holds color and opening fixed

The LLM always played Black from the standard initial position. Candidate order was randomized and scores and ranks were hidden, but one color and one opening family remain a narrow slice of chess. The inverse curve may be specific to this prompt, this candidate generator, this time control, or this provider snapshot. A protocol-specific regularity is still worth studying. It should not be renamed a model trait before it travels.

The experiment that can turn shape into effect

The decisive next step is a nested-set paired candidate-count crossover. For each frozen test position, one engine pass constructs a canonical top-10 parent set. The k=2, k=3, and k=5 conditions are strict nested subsets of that same parent. Every condition sees the same root position in a fresh context, under the same reasoning budget and action schema.

The design needs more than shared FENs:

  1. Pair every width within position. Run k in {2, 3, 5, 10} on every eligible position rather than assigning different positions to widths.
  2. Repeat presentation orders. Shuffle each displayed set with registered seeds, never expose engine ranks or scores, and balance candidate slots so a width effect cannot be a primacy effect in disguise.
  3. Use both colors and isolated source families. Balance side to move and cluster uncertainty by source game; keep normalized and color-mirrored families in one partition.
  4. Score the decision, not only survival. Record legality, chosen source rank, independent-evaluator WDL regret, catastrophic-blunder rate, response stability under order permutations, tokens, latency, and protocol failure.
  5. Freeze stopping and contrasts before outcomes. The primary adjacent contrasts are 2-3, 3-5, and 5-10, with source-game clustered uncertainty and a declared multiplicity rule. No adaptive anchor is needed for a position-level comparison.

Nested sets isolate a meaningful intervention: what happens when additional known candidates are appended to the same strong core. They do not, by themselves, explain why performance changes. A second decomposition must vary candidate provenance and distractor quality at fixed width: model-proposed moves, oracle-included sets, best-move-omitted sets, controlled legal distractors, and all-legal sets. That separates proposal coverage from selection load and tests whether “fewer is better” survives when quality is controlled.

The full-game ladder and paired position study then play complementary roles. The ladder measures end-to-end survival. The paired crossover measures the local consequence of widening one consideration set. Agreement between them would be far stronger than either alone; disagreement would reveal that trajectory dynamics, not local choice overload, produced the Elo curve.

What would change our mind

The comparison-overload interpretation should weaken if any of the following occurs:

  • the terminal DeepSeek freeze changes or erases the interim ordering;
  • paired nested sets show no adverse k=10 contrast after candidate identity, position, order, color, and budget are controlled;
  • the apparent width effect vanishes when candidate quality is held fixed;
  • order permutations explain most of the regret difference;
  • the effect appears only in full-history games and not in fresh-position decisions; or
  • prospective fixed-anchor, balanced-color games fail to reproduce the Luna ordering.

The opposite result would also be informative. If the same model chooses well from a nested top-2 core but degrades as matched distractors are added—and that effect survives order counterbalancing, both colors, independent evaluation, and source-game clustering—then “comparison load” becomes an empirical mechanism rather than a name for a surprising graph.

Evidence and reproduction

This essay is bound to the read-only ladder synthesis with analysis identity [private hash withheld]. Its exact synthesis cutoff is 2026-08-05T04:06:07.488774+00:00 UTC, the runtime-artifact cutoff recorded in that synthesis. The source JSON file SHA-256 is [private hash withheld]. DeepSeek values are the interim values in that artifact, not a live refresh and not a terminal freeze. The historical table and hashes above are intentionally not refreshed from later artifacts. The later terminal chronology and later terminal rating-uncertainty artifact are cross-links for chronology, not publication authority for this snapshot; release inventory reconciliation remains open.

uv run python [local path withheld]
uv run pytest -q [local path withheld] [local path withheld]
  [local path withheld]
Public evidence boundary

Reference destinations are classified from this public copy. Hash-only references return here; unavailable references retain a safe label but expose no private path or mutable artifact. External URLs are not verified by this build.

Public essay links
0
External URLs
0 (not verified)
In-page hash references
8
No public destination
0