The Shuffle That Looked Like a Signal
If a language model sees the same chess moves in a different order, does it choose differently?
The naive experiment is easy: show one list, shuffle it, ask twice, and count changes. It is also wrong. A model can change its answer even when the prompt is byte-for-byte identical. Without measuring that baseline instability, ordinary sampling noise masquerades as an ordering effect.
We ran the stronger version. For each of 48 frozen Black-to-move positions, Luna received the same five legal candidates in two visible orders. The second order was a complete cyclic derangement: no move remained in the same slot. Each order was repeated twice in a fresh context. The execution schedule was balanced between ABBA and BAAB blocks and frozen before the first response.
Cross-order decisions disagreed 42.7% of the time. That sounds dramatic until we look at the controls. The two identical A prompts disagreed on 39.6% of positions; the identical B prompts disagreed on 35.4%. After subtracting that same-order instability, the preregistered excess was only 5.2 percentage points. Its bootstrap interval ran from -2.1 to 13.5 points, and the within-position randomization reference was p=0.191.
So the careful result is not “order never matters.” The experiment remains compatible with a modest, potentially meaningful positive effect. It says we could not cleanly separate that effect from Luna's already substantial fresh-context variability with 48 positions and two repetitions per arm.
There was a tempting secondary pattern: the first visible slot was selected 58 times, compared with 38, 25, 29, and 42 for slots two through five. But this was not the primary controlled contrast, and candidate identity still matters. We record it as a follow-up signal, not a headline.
The run also caught a transport-boundary bug: one provider response returned
HTTP 200 with an undecodable body, and the frozen runner retried the identical
cell successfully. The locked v1 analyzer had treated any http_status event
as terminal. Amendment 1 froze corrected v1.1 retry semantics before outcome
analysis; the primary estimate stayed identical. The analysis amendment (In-page hash only)
keeps the original and corrected analyses together.
E025 does not tell us which move was best. It does not measure Elo, search, or chess vision. It tells us something narrower and useful: in this candidate-menu protocol, on development-exposed frozen Black-only positions, with two repetitions per order, heterogeneous provider backends, and one dated Luna alias cohort without an immutable provider snapshot ID, most apparent cross-order instability was already present when the order did not change. This is not a move-quality result; no independent evaluator scored the moves. The terminal aggregate is documented in the E025 report (In-page hash only), the public aggregate artifact (In-page hash only), and the public experiment page (In-page hash only).
Reference destinations are classified from this public copy. Hash-only references return here; unavailable references retain a safe label but expose no private path or mutable artifact. External URLs are not verified by this build.
- Public essay links
- 0
- External URLs
- 0 (not verified)
- In-page hash references
- 4
- No public destination
- 0