← All writing
Essay 27 / 27Frozen evidence

The Prompt Did Not Whisper

Showing Luna five unranked moves changed which move it submitted. That is a prompt-guidance result—not a chess-strength result.

After two failed transports, E028 finally measured the question E026 was built to ask.

The effect was not subtle. When Luna saw a fixed list of five legal moves, 92 of 96 returned actions belonged to that list: 95.83%. Given the same positions without the list, 2 of 96 actions belonged to it: 2.08%. The preregistered, position-weighted paired difference was 0.9375, or 93.75 percentage points. Its 10,000-draw position-bootstrap 95% interval was 0.8854 to 0.9792. The registered within-position randomization test gave a two-sided p-value of 0.0000099999.

On these positions, under this prompt and route, visible suggestions strongly guided Luna toward the displayed set.

That sentence is the result. It is intentionally narrower than “the model played better chess.”

What changed—and what did not

E028 used 48 development-exposed positions inherited from the earlier candidate-order study. Each position contributed four fresh calls: two visible and two board_only. The ABBA/BAAB assignment was frozen before the responses. Both conditions used the identical unrestricted submit_move tool, which accepted any move-shaped string. No schema forced the model to select one of the five suggestions.

Only the prompt differed. The visible condition displayed the same fixed set of five legal candidate moves. The board-only condition did not.

This is why the arm contrast identifies prompt guidance rather than menu compliance imposed by software. The model remained technically free to submit another move in both arms. The intervention changed the information in its context, not the legal action space exposed by the tool.

The estimand was equally weighted over positions. It was not a pooled fraction chosen after seeing the data. Each action counted one if it belonged to the frozen five-move set and zero otherwise. Illegal strings, malformed tool calls, refusals, and legal moves outside the set all counted zero. Provider failures would have remained missing, but none occurred.

The cohort that finally stayed intact

All 192 registered cells reached a terminal model action. There were 192 paid chat intents, zero chat retries, zero provider-error cells, and no unattempted cells. The runner reached the registered maximum of ten concurrent position blocks.

Metadata did not become instant. The terminal ledger records 737 read-only generation lookups: 31 cells resolved on the third scheduled lookup and 161 on the fourth. E028 did not treat that delay as a surprise to patch around. The absolute lookup schedule—0, 2, 5, 10, 20, 40, and 80 seconds after durable chat return—was frozen prospectively from the E027 incident. Lookup retries could not create a second chess answer, because chat retries remained forbidden.

Every accepted record resolved to the preregistered canonical model openai/gpt-5.6-luna-20260709, behind the requested alias openai/gpt-5.6-luna. Authoritative request-level costs summed to $0.1466976: $0.0578290 in the visible arm and $0.0888686 in the board-only arm. That cost is an accounting result for this exact cohort, not a general price estimate.

A large effect with a small interpretation

The strongest possible overclaim would be to equate candidate-set membership with chess quality. E028 never evaluated whether any listed move was best, whether listed moves were better than the model's alternatives, or whether following the list improved a game.

The five moves were fixed inputs. A move could be in the set and strategically poor. A board-only move could fall outside the set and be excellent. Because no quality score enters the primary outcome, the 93.75-point effect cannot be converted into centipawns, win probability, or Elo.

Nor does the experiment locate an internal mechanism. “Guidance” here is a causal description of behavior under a prompt intervention. It does not tell us whether Luna visually recognized a candidate, copied a salient token, reasoned about the list, searched more efficiently, or deferred to an apparent expert. Those mechanisms make different predictions and require new interventions.

The board-only membership rate is also not a general estimate of Luna's natural overlap with good candidate moves. The positions were development-exposed and the five-move sets came from an earlier experimental substrate. This was not a held-out sample from chess, and the list was not constructed to estimate an independent model policy.

What the result changes

Before E028, the candidate-set paradox had at least two live explanations. A small menu might improve performance because it narrows the tool's legal action space, or because seeing candidates changes the model's policy. E028 removes the first mechanism and finds a very large effect on set membership. Prompt information alone can dominate which move this Luna system submits.

That makes candidate presentation a first-class experimental variable. Future benchmarks cannot treat a list of moves as a neutral convenience. Candidate identity, order, source, explanation, score visibility, and distractor quality may all shape action independently of the model's underlying ability to generate or evaluate moves.

The next useful tests are mechanistic separations:

  • vary the fraction of strong, weak, and adversarial suggestions while keeping set size fixed;
  • compare bare moves with explanations from matched and mismatched positions;
  • measure whether guidance survives randomized notation and order;
  • score selected moves with a prospectively frozen independent evaluator; and
  • test held-out positions in both colors before connecting adherence to chess performance.

E028 establishes leverage. It does not establish benefit.

Evidence identity

The terminal result binds these frozen artifacts:

  • E028 preregistration, SHA-256 eba057087b90459b7d4902d1f43110f38952d7ef66ee00600882d2736c1b13f2;
  • E028 source and remediation audit, SHA-256 ed0faa246db8ac7d1cb7e210d0b29923da69e60b3d80edca105716daf4307922;
  • E028 manifest, SHA-256 13f053024f01fb08eab1d9927b4690dc848ea35658602c5d9c391a95ebfb401c;
  • E028 execution lock, SHA-256 091dabf059ad1bf3fb25e949c46d95421a1c745d73a8df947bdc4ab28a326f5c;
  • E028 completion, SHA-256 9bf6791a04c955d796fb27a7c91bb720d8b3fadb69cbb6bff5c5c7164ea093b9; and
  • E028 registered analysis, SHA-256 58e88291143c45f6b36c3ef7fe00f2683f53ec67a1b761a7ec218334ac8bf1ec.

The analysis's internal content identity is 68be55078665d9110bb77d18be77aea1e0a8d808b8fbea7fe842fdad163a4a6f. Its explicit boundary is the development-only causal effect of showing one fixed five-move suggestion set on Luna set adherence. It makes no move-quality, Elo, vision, search, candidate-generation, or general chess-strength claim.

Public-surface note

Operational paths, raw logs, credentials, mutable run metadata, and unpublished artifacts are intentionally excluded. Research references resolve to the versioned publication bundle rather than the host filesystem.