The Result That Is Still Zero
Historical pre-execution snapshot — 2026-08-05: E026 asked whether suggestions guide action. Its design and source audit were complete; its outcome was not yet known.
There is a satisfying kind of research progress that produces no number.
E026 is designed to ask a narrow causal question about Luna: if the model sees five unranked move suggestions, does it choose from that set more often than when it sees the same board without the suggestions? The action tool remains the same in both conditions. It accepts any move-shaped string; it does not constrain the model to the five visible moves.
That was the experiment as described in this pre-execution snapshot. At the
snapshot date, the project owner had frozen narrow budget consent for E026,
but that consent did not waive the remaining gates. The historical registry
fields bound to this snapshot recorded
preregistered_execution_no_go_pending_terminal_and_transport_gates,
execution_authorized as false, and paid_calls as 0; those are snapshot
facts, not a current-status line.
The current frozen authority records six paid chat intents, zero chat returns, and zero scientific outcomes (E026 transport incident freeze (In-page hash only)). E026 v1 therefore has no action, adherence rate, arm contrast, confidence interval, p-value, cost result, or chess result to report. Essay 26 (In-page hash only) records the terminal transport boundary.
The question that earned a protocol
Earlier candidate-menu experiments could not separate two mechanisms. A model
might choose a displayed move because the prompt guides its attention, or
because the tool schema makes every other action unavailable. E026 removes the
second mechanism. Both arms use the identical unrestricted submit_move
schema. Only the visible arm receives the five-move suggestion list.
For each of 48 frozen positions, the registered schedule contains four fresh
calls: two visible and two board_only. The position-level estimand is the
difference in candidate-set membership between those arms, averaged equally
over positions. Membership is intentionally blunt. A listed action counts as
one; another legal move, an illegal string, a malformed action, or a refusal
counts as zero. Provider and transport failures remain missing rather than
being converted into model failures.
The counterbalance is fixed before outcomes: each position uses ABBA or BAAB, with 24 positions assigned to each sequence. The maximum is 192 provider calls, with no provider-call retries. Neither an interim action nor an interim arm contrast may change the sample size or stopping rule.
This design was intended to estimate prompt-level guidance on this exposed development set. It does not estimate move quality, candidate quality, Elo, internal attention, vision, search, candidate generation, or the value of showing fewer moves.
A qualified source is not an authorized experiment
The source-reuse audit reached a qualified GO for one limited purpose. E026 may reuse the complete 48-position E025 manifest because those positions and their five-move sets were frozen before E025 responses. The reusable source contains FENs, legal candidate identities, lineage, and prospective order assignments; it does not contain E025 selected moves or response outcomes.
That source qualification is not permission to send request one.
The audit also records why the study remains developmental. E025 has exposed the positions, and its published result motivated this measurement question. E026 therefore cannot be called a held-out confirmation. Its positions cannot later enter a confirmatory held-out pool. The inherited E000 candidate lists are not the stable nested candidate products planned for E020, E021, and E024, so E026 cannot answer their candidate-quality or candidate-width questions.
The source audit is valid only while hard exclusions remain intact. E026 may not read E025 responses to select positions, score moves with prior results or a new engine, change its sample from observed outcomes, or substitute for the larger generation, selection, and candidate-width experiments. Remove any one of those restrictions and the source disposition becomes NO-GO.
The pre-execution gates
In this historical snapshot, the preregistration required every execution gate to pass before the first paid request. The exact design and implementation bytes had to be locked. The source manifest and its pre-response lock had to be rebound. The E001 terminal freeze had to be verified and no ladder screen or process could remain active. Luna's selected chat route had to have a genuine qualification. A separate, content-hashed E026 paid authorization and an immutable execution lock had to exist. The authenticated budget check had to pass, and every irreversible intent and returned exchange had to enter the paid-call journal.
At the snapshot date, the narrow owner authorization existed, so one gate was
satisfied. The snapshot registry still recorded execution_authorized: false.
A well-written preregistration and a qualified source audit were necessary
artifacts, not substitutes for authorization.
What this snapshot did not claim
The six paid E026 v1 intents do not measure a zero guidance effect. They produced no returned exchanges from which an adherence contrast could be estimated. The correct current boundary is no scientific outcome, not “zero effect.”
E026 v1 is frozen: no resume, retry, or added v1 intent is permitted. Any future attempt would be a separately versioned E026 v2 with separately qualified live transport. See Essay 26 (In-page hash only) for the terminal transport boundary and Essay 27 (In-page hash only) for the later E028 adherence result.
Reference destinations are classified from this public copy. Hash-only references return here; unavailable references retain a safe label but expose no private path or mutable artifact. External URLs are not verified by this build.
- Public essay links
- 0
- External URLs
- 0 (not verified)
- In-page hash references
- 4
- No public destination
- 0