Experiment A105 / local preflight only

A controlled test of
root-information access.

Does redundant access to the root board or complete legal-move list change system regret when the underlying chess state, action universe, response contract, and request length stay fixed?

2 × 2 treatment factors4 treatment cells2 panel orders0 observed model outcomes

01 / Treatment cells

Four cells, one retained state.

The factorial design varies board-content access and legal-list access. Each position is rendered in both registered panel orders, while the state, action universe, response contract, and request length remain common.

Board content →Legal-list content →
Selected treatment cell

B0 · L0 / Masked + masked

Board: maskedLegal list: masked

Both treatment panels preserve their shape but replace chess identities with equal-length masks.

Board panel64 equal-length masksRegistered board-content span
Legal-action panelequal-length move masksRegistered legal-content span

02 / Paired design

What every cell retains.

All arms contain the same chess state and legal-action universe. Pairing unit: the same position appears in both panel orders, and the analysis averages those orders at each position before inference.

  1. 01Full root FEN and side-to-move task
  2. 02Complete legal-action universe
  3. 03Identical submit-move response schema
  4. 04Identical request parameters and byte length
  5. 05Identical panel labels, delimiters, slots, and token lengths
  6. 06The same position in both panel orders

03 / Renderer contract

Registered spans only.

Shape-matched masks make the treatment auditable at byte level. A difference outside the two registered content spans is an automatic no-go.

Common requestBoard spanCommonLegal spansCommon response contract
01

Board-content span

Only the 64 registered square-token bytes may switch between piece symbols and one-byte masks.

02

Legal-action span

Only registered move-token bytes may switch between sorted UCI actions and equal-length masks.

03

Everything else

Any difference outside the two registered spans is an automatic no-go.

16fixture positions
128rendered cells
32matched blocks
0contract failures

The local renderer measured 2,048 board-treatment bytes and 2,104 legal-treatment bytes inside 110,640 canonical request bytes. These are compiler measurements, not model observations.

04 / Analysis contract

Primary inference, then sensitivity.

The synthetic preflight fixes failure handling, dependence, multiplicity, and reporting rules before any future model outcome. The current snapshot contains no observed model outcome.

256synthetic analysis cells
0observed model outcomes
0provider calls
0network calls

Engine calls in this snapshot: 0. The zero outcome and call counts describe preflight status; they are not a model result.

Step 01 / 06

Freeze and blind

Write only frozen condition codes to outcome rows. Hold the code key separately until the complete cell grid and analysis digest are frozen.

PRIMARY / ESTIMAND

Provider-complete pairs

ITT (intention-to-treat) keeps each randomized assignment; the primary readout uses provider-complete pairs.

PRIMARY / DEPENDENCE

Position-level pairing

Average the two panel orders at each position and weight positions equally. Use source-game cluster bootstrap (resample whole source-game clusters) only for intervals and source-game sign-flip randomization (reverse each cluster's paired difference) for primary p-values.

PRIMARY / MULTIPLICITY

Four future tests

Apply Holm step-down correction (false-positive safeguard) across exactly four future tests: board and legal-list effects, reported separately for each model.

SENSITIVITY / EXECUTION

Maximal regret 1.0

A model execution failure receives maximal regret 1.0. It stays in the randomized comparison instead of disappearing as a convenient exclusion.

SENSITIVITY / MISSINGNESS

Provider attrition

An exhausted provider failure remains missing. Report complete-pair inference plus all-position best/worst attrition bounds.

SENSITIVITY / REPORTING

Null and adverse cases

Report null and adverse outcomes, successful-only sensitivity, attrition bounds, route and order diagnostics, and material sign reversals alongside any gain.

Intervals2,000 source-game bootstrap replicates
Primary p-valuesSource-game sign-flip randomization
Exact sign-flip limit16 source-game clusters
Above exact limit20,000 Monte Carlo replicates (random assignment samples)
MultiplicityHolm across four tests

05 / System checks

Local checks have a boundary.

Provider-free stress tests expose transport and client edges without promoting local captures into provider facts.

R0S / opaque-mask sensitivity audit

Mask sensitivity

4 opaque families preserved the byte/span contract across 512 cells. Equal bytes and registered spans do not imply equal tokenization, salience, or semantic difficulty. Exact offline tokenizers were unavailable, so proxy counts were prohibited.

R2 / transport capture check

Transport capture

256 captured requests crossed 2 frozen profiles with zero registered drift. The repository adapter reached an in-memory sink without arm-dependent route, header, tool, parameter, or out-of-span body drift. This is not proof of AutoGen, HTTP, proxy, router, or provider-received bytes.

R3 / installed-client boundary check

Client boundary: chat pass, Responses NO-GO

256 installed-stack requests were arm-invariant at the injected pre-network HTTP boundary. Chat preserved required submit_move semantics; Responses did not. The installed AutoGen Responses adapter rewrote required submit_move tool choice to auto before HTTP serialization. Captured at an injected httpx transport after installed AutoGen normalization and OpenAI SDK serialization, before network transport. This is not proof of provider-received bytes or provider compatibility; paid execution remains unauthorized.

AutoGen 0.11.2 · OpenAI Python 2.30.0 · httpx 0.28.1
R3A / isolated remediation candidate

Isolated remediation candidate—not a repaired history

128 remediated Responses captures formed 32 arm-invariant blocks and preserved required submit_move. An isolated wrapper changed only the installed adapter's tool_choice=auto output to required submit_move before SDK serialization. R3 remains the observed installed-stack NO-GO; this candidate is not in the live harness, proves no provider compatibility or provider-received bytes, and does not authorize paid execution.

R3 status preserved: partial_chat_pass_responses_no_go

06 / Claim boundary

Separate local measurement from causal evidence.

Measured locallyThe local request compiler, stipulated synthetic analysis contract, capture-only adapter, installed chat client boundary, and isolated Responses remediation candidate pass their respective version-bound local gates.

Causally excludedR3's installed Responses path remains the observed NO-GO; R3A is isolated and not integrated into the live harness. No provider-received bytes, provider compatibility, provider behavior, model effect, chess-quality effect, power result, final-position result, paid execution, or causal benefit has been observed.

Every arm still contains the same full FEN and action universe. A105 tests redundant representation and access—not whether the model can play chess without state.