← All writing
Essay 25 / 25Protocol / design

The Result That Is Still Zero

E026 asks whether suggestions guide action. It has a design, a source audit, and zero outcomes.

There is a satisfying kind of research progress that produces no number.

E026 is designed to ask a narrow causal question about Luna: if the model sees five unranked move suggestions, does it choose from that set more often than when it sees the same board without the suggestions? The action tool remains the same in both conditions. It accepts any move-shaped string; it does not constrain the model to the five visible moves.

That is the experiment. It has not run.

The project owner has now frozen narrow budget consent for E026, but that consent does not waive the remaining gates. The current registry status is preregistered_execution_no_go_pending_terminal_and_transport_gates. execution_authorized is false, paid_calls is 0, and the claim field is none_until_terminal_analysis. There is no E026 action, adherence rate, arm contrast, confidence interval, p-value, cost result, or chess result to report.

The question that earned a protocol

Earlier candidate-menu experiments could not separate two mechanisms. A model might choose a displayed move because the prompt guides its attention, or because the tool schema makes every other action unavailable. E026 removes the second mechanism. Both arms use the identical unrestricted submit_move schema. Only the visible arm receives the five-move suggestion list.

For each of 48 frozen positions, the registered schedule contains four fresh calls: two visible and two board_only. The position-level estimand is the difference in candidate-set membership between those arms, averaged equally over positions. Membership is intentionally blunt. A listed action counts as one; another legal move, an illegal string, a malformed action, or a refusal counts as zero. Provider and transport failures remain missing rather than being converted into model failures.

The counterbalance is fixed before outcomes: each position uses ABBA or BAAB, with 24 positions assigned to each sequence. The maximum is 192 provider calls, with no provider-call retries. Neither an interim action nor an interim arm contrast may change the sample size or stopping rule.

This design estimates prompt-level guidance on this exposed development set. It does not estimate move quality, candidate quality, Elo, internal attention, vision, search, candidate generation, or the value of showing fewer moves.

A qualified source is not an authorized experiment

The source-reuse audit reached a qualified GO for one limited purpose. E026 may reuse the complete 48-position E025 manifest because those positions and their five-move sets were frozen before E025 responses. The reusable source contains FENs, legal candidate identities, lineage, and prospective order assignments; it does not contain E025 selected moves or response outcomes.

That source qualification is not permission to send request one.

The audit also records why the study remains developmental. E025 has exposed the positions, and its published result motivated this measurement question. E026 therefore cannot be called a held-out confirmation. Its positions cannot later enter a confirmatory held-out pool. The inherited E000 candidate lists are not the stable nested candidate products planned for E020, E021, and E024, so E026 cannot answer their candidate-quality or candidate-width questions.

The source audit is valid only while hard exclusions remain intact. E026 may not read E025 responses to select positions, score moves with prior results or a new engine, change its sample from observed outcomes, or substitute for the larger generation, selection, and candidate-width experiments. Remove any one of those restrictions and the source disposition becomes NO-GO.

Why execution is still NO-GO

The preregistration requires every execution gate to pass before the first paid request. The exact design and implementation bytes must be locked. The source manifest and its pre-response lock must be rebound. The E001 terminal freeze must be verified and no ladder screen or process may remain active. Luna's selected chat route must have a genuine qualification. A separate, content-hashed E026 paid authorization and an immutable execution lock must exist. The authenticated budget check must pass, and every irreversible intent and returned exchange must enter the paid-call journal.

The narrow owner authorization now exists, so one gate is satisfied. The current registry does not claim the remaining gates have passed. It says execution_authorized: false. That is the controlling public fact. A well-written preregistration and a qualified source audit are necessary artifacts, not substitutes for authorization.

What a future result could say

If the campaign is eventually authorized and terminalized, a positive paired adherence difference would support a narrow statement: on these 48 exposed positions, under this route and prompt, visible suggestions guided Luna toward the displayed set even though the action schema allowed any string. A zero or negative difference would falsify that narrow guidance effect under the same conditions.

Neither result would show that the listed moves were good. Neither would show that the model understood why a move appeared, generated it independently, or played stronger chess because it saw it. E020 and E021 remain the registered generation-and-selection tests. E024 remains the candidate-width test.

Until a terminal analysis exists, the correct E026 result is not “pending but probably positive,” nor “zero effect.” It is no outcome.

Evidence identity

This status draft is bounded to three source artifacts:

  • E026 preregistration, SHA-256 dfa532a774ee7b36bd38f4340212c8308caff57409a0b93028313c9130e4b49e;
  • E026 source-reuse audit, SHA-256 d5f3459fb0a07b571cbf31e145e746eb0bfa5d87d51a7f9d92ce72b3d80cfb61; and
  • E026 owner authorization, SHA-256 6c1f163287dce42d8b2a56857c9fb5b0755d0a53a0d308f702f7f114ea043f33; and
  • experiment registry, SHA-256 45b773eb668d45278fa54916ea11f71c1797048dcb1b890acf4c8d20ff164bfc.

If any identity changes, this draft must be re-audited before it is described as current. It does not bind a future execution manifest or terminal result, because neither exists here.

Public-surface note

Operational paths, raw logs, credentials, mutable run metadata, and unpublished artifacts are intentionally excluded. Research references resolve to the versioned publication bundle rather than the host filesystem.