← All writing
Essay 26 / 30Frozen evidence

The Result That Arrived Late

E026 and E027 produced no scientific outcome. They produced the transport contract that made E028 interpretable.

We tried to measure a prompt-guidance effect twice, paid for sixteen chat intents, and refused to read either attempt as an experiment.

That refusal was not ceremony. It was the result.

E026 failed before a durable model response entered the evidence record. E027 returned ten chats concurrently, but the metadata needed to attest those responses was not visible when the registered runner looked for it. In both cases the tempting move was to repair the machinery around the data already in flight and keep going. In both cases that would have changed the measurement after seeing how the provider behaved.

So the laboratory froze each failure and started a new version.

E026: six intents, zero returned exchanges

Essay 25 records E026's preregistered question and unrestricted-tool design; its v1 campaign never became an outcome dataset.

Its immutable incident freeze records six registered chat intents. Five ended in APIConnectionError; one became an ambiguous intent with no durable return. There were zero durable chat returns, zero generation-metadata records, zero provider-call retries, and zero scientific outcomes. The evidence localized the problem only to the live adapter or downstream network boundary. It could not recover an exact leaf cause.

The correct E026 estimate is therefore not zero. It is not estimable.

Coding a transport failure as a model failure would have made an arm look worse for a reason unrelated to the model's action. Repeating an ambiguous paid intent could have created an unknown duplicate. Resuming after changing the route would have mixed two transports inside one preregistered cohort. The freeze prohibited all three shortcuts.

E027: ten returned chats, ten missing attestations

E027 was a fresh operational qualification, not a continuation of E026. It used direct concurrent requests and admitted ten position blocks at once. The question was deliberately operational: could the 192-cell design run with ten concurrent Luna blocks without depending on the slow DeepSeek ladder?

On that narrow question, concurrency worked. The first frontier admitted ten blocks, sent ten chat intents, and received ten chat returns. The maximum observed concurrency was ten. There were no chat retries.

Then every immediate generation-metadata lookup returned HTTP 404.

The metadata was not decorative. It was the authoritative evidence for cost, provider identity, and the resolved model behind the stable alias. Without it, the runner could not prove that a returned action belonged to the registered route. The first frontier stopped the campaign exactly as registered; the remaining 182 cells were never admitted, and the ten response actions stayed quarantined from scientific analysis.

After E027's terminal freeze, one delayed read-only forensic lookup sampled one generation record. It made no chat call and read no action. The record returned HTTP 200, included cost and provider identity, and resolved the requested alias openai/gpt-5.6-luna to the dated model openai/gpt-5.6-luna-20260709. That established two facts and no more: metadata could arrive after the chat response, and exact equality between a stable alias and a canonical model string was the wrong identity test. It did not authorize reopening E027 or reading its chess actions.

Why the failed records stayed failed

The chats existed, and a human could have inspected them; a small script change could have waited longer or accepted the dated suffix. But either choice would have used the failure itself to rewrite the inclusion rules for the same cohort, turning a prospective protocol retrospective and erasing why the protocol changed.

Instead, E027 was terminalized with ten provider-missing attempt records and no scientific estimate. The delayed lookup was stored as operational evidence in a separate incident artifact. E028 then preregistered the correction before request one: accept one exact canonical resolved model, poll metadata on a fixed absolute schedule, retry only the read-only lookup, never retry a chat, and stop admission if attestation is exhausted or inconsistent.

This is what versioning an experiment is for. E026 proved that local request serialization was not live-route qualification. E027 proved that a successful chat return was not yet a fully attested observation. Neither proved anything about whether visible suggestions guide Luna's move.

The scientific lesson in an operational failure

An LLM experiment has more failure surfaces than its statistical model admits. There is the prompt, model, action parser, network, routing layer, delayed billing metadata, model alias resolution, journaling, and crash recovery. If a surface can change which observations enter an arm, it belongs inside the experimental protocol.

The most valuable artifact from E026 and E027 was a more exact definition of an observation:

  1. one durable registered chat intent;
  2. one durable returned exchange;
  3. one matching generation record;
  4. the accepted canonical model identity;
  5. authoritative cost and provider metadata; and
  6. one terminal, hash-chained event.

Only after all six existed could an action enter the E028 analysis.

That standard costs time. It also prevents a more expensive mistake: publishing a clean-looking effect whose inclusion rule was invented while the data arrived.

Public evidence boundary

Reference destinations are classified from this public copy. Hash-only references return here; unavailable references retain a safe label but expose no private path or mutable artifact. External URLs are not verified by this build.

Public essay links
0
External URLs
0 (not verified)
In-page hash references
1
No public destination
0