Research · published whether it suits us or not

Preregistration — Stage 1b

Why there is a Stage 1b at all

Two things were learned from Stage 1, one a defect in our own instrument and one a defect in how the claim was operationalised. Both are stated here before the new run, and the new run uses a fresh seed so that nothing is confirmed on the data that suggested it.

D1 — an oracle-validity defect that favoured a competitor

The Stage 1 game client used a fixed interpolation delay of two ticks. A client that never adapts must extrapolate further into the future exactly in proportion to the path's mean delay, so its damage is close to a monotone function of mean one-way delay by construction. owd_mean did not win a fair contest; it was handed one.

Shipped clients adapt their interpolation buffer to observed conditions. Stage 1b therefore runs both clients and reports both:

This change is expected to help the hypothesis, because absorbing constant delay leaves variation as what damages the client. That is stated here, in advance, precisely because it is the kind of change that is otherwise indistinguishable from tuning until the answer comes out right. The justification is realism, and the fixed client is retained so the effect of the choice is visible rather than hidden.

D2 — the claim was operationalised as a point when the product ships a curve

Bullettime's actual output is the deadline survival curve, not a miss rate at one guessed deadline. Stage 1 tested a point on that curve. Two better-posed forms are added:

Fixed in advance

Falsification, stated before the result

Stage 1b does not support the claim if, on the adaptive client's RMS positional error, the 95% CI for |ρ(C3)| − |ρ(P_best)| includes zero or lies below it in the conditional analysis. The conditional analysis is the decisive one, because the unconditional one can be won by any predictor that happens to track absolute delay.

If Stage 1b also fails, the programme stops and the marketing page and the PRD are changed to say the mechanism is real, the comparative claim was tested twice and did not hold, and no subjective panel will be run.

Harness, seeds and raw JSON live beside these documents in the repository. Every figure here comes from simulated impairment; no human subjects were involved and no perception claim is made. Back to Bullettime.