The Prompt Did Not Whisper
Showing Luna five unranked moves increased adherence to the displayed set. That is a prompt-guidance result—not a chess-strength result.
E026's live transport aborted before any chat return; E027's chat transport succeeded for ten returns, but metadata attestation failed on the immediate lookup. E028 then ran as a fresh response cohort with zero E027 actions or outcomes imported and measured the question E026 was built to ask.
The effect was not subtle. When Luna saw a fixed list of five legal moves, 92 of 96 returned actions belonged to that list: 95.83%. Given the same 48-position set without the list, 2 of 96 actions belonged to it: 2.08%. The preregistered, position-weighted paired difference was 0.9375, or 93.75 percentage points. Its 10,000-draw position-bootstrap 95% interval was 0.8854 to 0.9792. The registered within-position randomization test gave a two-sided p-value of 0.0000099999.
On these positions, under this prompt and route, visible suggestions strongly guided Luna toward the displayed set.
That sentence is the result. It is intentionally narrower than “the model played better chess.”
What changed—and what did not
E028 used 48 development-exposed positions and the unrestricted four-cell
design described in Essay 25 (In-page hash only), but as a
fresh response cohort: zero E027 actions or outcomes were imported. Each
position contributed four fresh calls—two visible and two board_only—with
the ABBA/BAAB assignment frozen before responses. Both conditions used the
identical unrestricted submit_move tool, which accepted any move-shaped
string; no schema forced the model to select one of the five suggestions.
Only the prompt differed. The visible condition displayed the position-specific fixed set of five legal candidate moves. The board-only condition did not.
This is why the arm contrast identifies prompt guidance rather than menu compliance imposed by software. The model remained technically free to submit another move in both arms. The intervention changed the information in its context, not the legal action space exposed by the tool.
The estimand was equally weighted over positions. It was not a pooled fraction chosen after seeing the data. Each action counted one if it belonged to the frozen five-move set and zero otherwise. Illegal strings, malformed tool calls, refusals, and legal moves outside the set all counted zero. Provider failures would have remained missing, but none occurred.
The cohort that finally stayed intact
All 192 registered cells reached a terminal model action. There were 192 paid chat intents, zero chat retries, zero provider-error cells, and no unattempted cells. The runner reached the registered maximum of ten concurrent position blocks.
Metadata did not become instant. The terminal ledger records 737 read-only generation lookups: 31 cells resolved on the third scheduled lookup and 161 on the fourth. E028 did not treat that delay as a surprise to patch around. The absolute lookup schedule—0, 2, 5, 10, 20, 40, and 80 seconds after durable chat return—was frozen prospectively from the E027 incident. Lookup retries could not create a second chess answer, because chat retries remained forbidden.
Every accepted record resolved to the preregistered canonical model
openai/gpt-5.6-luna-20260709, behind the requested alias
openai/gpt-5.6-luna. Authoritative request-level costs summed to $0.1466976:
$0.0578290 in the visible arm and $0.0888686 in the board-only arm. That cost is
an accounting result for this exact cohort, not a general price estimate.
A large effect with a small interpretation
The strongest possible overclaim would be to equate candidate-set membership with chess quality. E028 never evaluated whether any listed move was best, whether listed moves were better than the model's alternatives, or whether following the list improved a game.
The five moves were fixed inputs. A move could be in the set and strategically poor. A board-only move could fall outside the set and be excellent. Because no quality score enters the primary outcome, the 93.75-point effect cannot be converted into centipawns, win probability, or Elo.
Nor does the experiment locate an internal mechanism. “Guidance” here is a causal description of behavior under a prompt intervention. It does not tell us whether Luna visually recognized a candidate, copied a salient token, reasoned about the list, searched more efficiently, or deferred to an apparent expert. Those mechanisms make different predictions and require new interventions.
The board-only membership rate is also not a general estimate of Luna's natural overlap with good candidate moves. The positions were development-exposed and the five-move sets came from an earlier experimental substrate. This was not a held-out sample from chess, and the list was not constructed to estimate an independent model policy.
What the result changes
Before E028, the candidate-set paradox had at least two live explanations. A small menu might improve performance because it narrows the tool's legal action space, or because seeing candidates changes the model's policy. E028 held the action space unrestricted, so its contrast measured prompt-level adherence to the displayed set rather than compliance imposed by the tool. It did not establish a prior performance mechanism or show that candidate presentation improved play; its causal result is the change in set membership on this development cohort.
That makes candidate presentation a first-class experimental variable. Future benchmarks cannot treat a list of moves as a neutral convenience. Candidate identity, order, source, explanation, score visibility, and distractor quality may all shape action independently of the model's underlying ability to generate or evaluate moves.
Next tests—not findings
The following are next tests, not findings from E028:
- vary the fraction of strong, weak, and adversarial suggestions while keeping set size fixed;
- compare bare moves with explanations from matched and mismatched positions;
- measure whether guidance survives randomized notation and order;
- score selected moves with a prospectively frozen independent evaluator; and
- test held-out positions in both colors before connecting adherence to chess performance.
E028 establishes leverage. It does not establish benefit.
Reference destinations are classified from this public copy. Hash-only references return here; unavailable references retain a safe label but expose no private path or mutable artifact. External URLs are not verified by this build.
- Public essay links
- 0
- External URLs
- 0 (not verified)
- In-page hash references
- 1
- No public destination
- 0