Experiment E025 / terminal result / not clearly supported

The shuffle moved.
The excess stayed uncertain.

With randomized candidate order (the same legal moves shown in different slots), and after controlling for repeated-call instability (the model can vary when order stays the same), the preregistered (specified before seeing outcomes) excess was +5.2 pp. Its 95% position-bootstrap interval (a resampling-based uncertainty range) was −2.1 pp to +13.5 pp; it crosses zero.

48 positions192 registered calls5 legal candidates per positioncorrected primary v1.1
Preregistered questionDoes visible order change Luna's selected action beyond same-order stochastic instability when candidate identities are fixed?
Denominator48 positions × 4 calls = 192 registered calls

2 fresh calls for each of two visible orders at every position.

Primary estimand (the target quantity this analysis is designed to estimate)Excess action disagreement

Cross-order disagreement minus the mean of the two same-order baselines (repeat-call variation with order held fixed).

Null value (no excess over baseline)0.0 pp

Cross-order disagreement equals the baseline-adjusted expectation.

Raw cross-order42.7%

Action disagreement when the same candidates changed visible order.

Same order A39.6%

Baseline disagreement between identical-order calls.

Same order B35.4%

The second identical-order baseline.

Adjusted excess+5.2 pp

95% interval −2.1 pp to +13.5 pp; randomization p-value (how unusual this result is under the no-excess shuffle)=0.191.

Candidate-order disagreement compared with repeated identical-order disagreement; the adjusted interval crosses zero.
The interval crosses zero: +5.2 pp with a 95% position-bootstrap interval of −2.1 pp to +13.5 pp. Conclusion: excess order sensitivity was not clearly supported.

Read the null correctly

Zero is a control result.
Not a universal answer.

The 42.7% raw cross-order disagreement is not itself an order effect. Luna also disagreed 39.6% and 35.4% of the time under repeated identical orders. The adjusted estimate is +5.2 pp, with an interval of −2.1 pp to +13.5 pp.

Zero is the null value, not the point estimate: a zero estimate would mean that changing visible order adds no excess disagreement beyond those same-order baselines for this estimand. It does not mean Luna was stable, establish a universal no-order-effect claim, or rule out a positive effect. E025 also does not measure move quality, Elo, search, or vision.

Evidence boundary. Action stability for one dated Luna alias cohort on development-exposed Black positions; not move quality, Elo, search, or a universal null. Open C024 evidence boundary

Why the null mattersThe 42.7% raw contrast was not enough to call order causal.
What remains openThe interval still permits a positive effect up to +13.5 pp.
Analysis correctionE025-analysis-amendment-1

one retryable undecodable response is no longer a terminal unknown failure. The corrected primary result is disclosed here rather than silently replacing the earlier analysis.

Limits carried into the page

  • interval includes zero and positive effects up to about 13.5 percentage points
  • two replicates per order
  • development-exposed Black-only positions
  • one dated Luna alias cohort without immutable provider snapshot ID
  • provider backend heterogeneity
  • no move-quality evaluator
  • protocol-specific
Read the laboratory note