Research · published whether it suits us or not

Stage 1 results — the central claim, tested twice, and what happened

Date: 26 August 2026 · Harness: research/lab/ · Raw: results/stage1.json, results/stage1b.json, results/stage1b-exploratory.json Preregistered before running: PREREGISTRATION.md, PREREGISTRATION-1b.md


The short version

The mechanism Bullettime is built on is not in doubt and is already standards consensus. The comparative claim the product makes on top of it — that a deadline-miss rate predicts damage better than the conventional metrics — was tested twice under preregistration and failed both times in the specific forms the product currently ships.

One form survives, and it is the least distinctive one: effective loss at a fixed, generous buffer depth, which is close to what ITU-T G.107 has done since 2015.

The part of the positioning that failed hardest is the part the marketing page leans on most: "the deadline is different for every application."


1 · What the literature actually says

Two independent research runs (reports/), one Claude-family and one Perplexity, agree:

So the PRD's sentence needs sharpening rather than deleting. It is not that nobody has connected lateness to quality; the standards did that twenty years ago. It is that nobody has run the head-to-head.

2 · What we ran

1,280 simulated sessions per stage over a 64-cell design (4 loss severities × 4 delay severities × {Bernoulli, Gilbert-Elliott} × {sustained queue, heavy-tailed spikes}), 30 s each. Loss severity and delay severity are drawn independently, which is what populates the off-diagonal cells where the predictors can be told apart at all.

Impairment is a time-varying process; a 50 pps probe lane and the application stream sample it separately, as they do in the product. Every predictor is computed from the probe lane only, using RFC 5481 PDV relative to the sample minimum, so no synchronised clock is assumed.

The oracle contains no deadline. Damage is the RMS positional error, in metres, between what a game client drew and where the entity authoritatively was, plus a shot-miss count. Delay autocorrelates (AR(1)) specifically to avoid the artefact both reports flag as the field's worst: tc-netem drawing delay independently per packet manufactures reordering no real FIFO queue produces. Measured reordering in our traces: 2.4%.

3 · Stage 1 — the pre-registered primary failed

predictorSpearman ρ vs positional error
owd_mean+0.769
owd_p95+0.744
deadline_miss_app (candidate, 16.7 ms)+0.675
pdv_p99+0.584
loss_pct+0.299
jitter_rfc3550−0.078

|ρ(candidate)| − |ρ(owd_mean)| = −0.094, 95% CI [−0.129, −0.061]. The CI excludes zero and lies below it. H1 not supported.

4 · Stage 1b — we found a defect in our own instrument, fixed it, and it failed again

Stage 1's client used a fixed interpolation delay, so it had to extrapolate further in proportion to mean delay. Damage was close to a monotone function of owd_mean by construction. owd_mean had not won a fair contest.

Stage 1b was preregistered before running, on a fresh seed, with an adaptive client (what shipped clients do), a better-posed candidate (dmc_auc, the survival-curve integral — the product's actual figure as a scalar), and a conditional analysis controlling for owd_mean, because every speed test already reports mean latency and the product's claim has only ever been about the residual.

Stating it in advance: that change was expected to help the hypothesis. It did not.

candidate dmc_aucbest conventionaldifference [95% CI]
adaptive client, RMS error, conditional+0.361loss_pct +0.655−0.295 [−0.337, −0.253]
adaptive client, shot misses, conditional+0.25loss_pct +0.42−0.172 [−0.209, −0.136]
fixed client, RMS error, conditional+0.47loss_pct +0.49−0.022 [−0.057, +0.013]
voice (circular, discounted)+0.53loss_pct +0.68−0.152 [−0.204, −0.100]

Not supported on every outcome. The pre-registered falsification criterion is met.

The mechanism behind that is interpretable and worth understanding: once a client adapts its buffer, delay variation is converted into added latency, and what remains unfixable is a packet that never arrived. Under adaptation, conventional loss dominates. That is not a bug in the experiment; it is what jitter buffers are for.

5 · What survived, exploratorily, and what it costs the positioning

Post-hoc and requiring its own confirmation: one deadline form is competitive.

metricρpartial ρ (controlling owd_mean)
deadline_miss_100 — effective loss at a fixed 100 ms+0.716+0.675
loss_pct+0.580+0.655
dmc_auc (the pre-registered candidate)+0.546+0.361
deadline_miss_app — at the application's own 16.7 ms tick+0.414+0.137

deadline_miss_100 beats loss_pct by +0.137 unconditionally [+0.109, +0.165] and by +0.020 conditionally [+0.003, +0.037], p = 0.021. Real, but small, and marginal once latency is controlled.

This is the finding that matters commercially. The form that survives is effective loss at a fixed generous buffer — essentially G.107's Peff, which is prior art. The form that fails hardest is the application-tuned deadline, partial ρ +0.137 against +0.675. The page currently says "the one metric in the set whose threshold genuinely varies by application". On this evidence that is the weakest part of the claim, not the strongest.

6 · What we are not entitled to say

7 · What this changes, today

  1. The PRD's open-question wording is now wrong in a specific way and should be replaced with the sharper statement in §1: the mechanism is standards consensus; the head-to-head is what is missing.
  2. The per-application deadline claim should come off the marketing page as a headline until it has evidence, or be restated as the mechanism it is rather than the advantage it is presented as.
  3. The survival curve remains defensible as a presentation, and its integral is not a good scalar predictor. Those are different claims and only the first is being made.
  4. No subjective panel should be commissioned yet. Stage 1 exists so that this decision is cheap, and it has just paid for itself.

Stages 1c and 1d — the challenge was right, and it found a better metric than either of us proposed

Added 26 August 2026, after the Stage 1b write-up was challenged on mechanistic grounds: late packets, depending on how late, must damage netcode. That challenge was correct, my Stage 1/1b operationalisation was wrong, and following it up changed the product recommendation.

A · The mechanism, measured at the unit where it operates

Stages 1 and 1b correlated session-level scalars against session-level damage. The mechanism does not operate at session level. Measuring it per render frame (95,733 frames, 40 sessions, adaptive client):

staleness of the client's informationframesmedian errorp95 error
0 (interpolating normally)82,5710.001 m0.002 m
10–25 ms2,3690.009 m0.019 m
25–50 ms7500.022 m0.066 m
50–100 ms3790.097 m0.267 m
100–200 ms4380.356 m0.987 m
>200 ms5911.961 m11.220 m

Spearman(staleness, error) among stale frames = +0.950. The relationship is nearly deterministic and graded over three orders of magnitude. 86% of frames take no damage at all.

Nothing in Stage 1 or 1b contradicted this and nothing could have: the oracle has no damage source other than stale or missing information. The earlier stages tested whether one summary statistic out-ranked others across sessions, which is a different question that I reported too broadly.

B · Why a miss rate was the wrong instrument

A rate thresholds lateness into a binary count, so it scores a packet 2 ms past the tick the same as one 400 ms past it, while the table above shows the damage spanning 0.001 m to 1.961 m across exactly that range. The tail carries the harm and a rate discards the tail's magnitude.

C · Stage 1c — magnitude helps, but does not win on its own

96 cells (a third delay regime added: spike, low mean delay with rare severe excursions, the regime the earlier grids under-sampled), 1,344 sessions, fresh seed. Partial ρ controlling for mean delay. Pre-registered H2: a magnitude-weighted lateness statistic beats conventional loss.

late_excess_mean +0.697 against loss_pct +0.690: difference +0.007, CI [−0.006, +0.021], Holm p = 0.27. H2 not supported. Magnitude closes almost all of the gap the rate form lost, and does not clear it.

D · Stage 1d — what actually wins, confirmed on a fresh seed

Stage 1c's top predictor was neither family under test. It was burst_max_ms, the longest continuous hole in the data, in milliseconds — a pre-specified predictor in 1c but not its candidate, so it required confirmation on data that had not suggested it. Seed 20260829:

adaptive clientfixed client
burst_max_ms+0.820+0.690
late_excess_mean+0.693+0.629
loss_pct+0.692+0.540
deadline_miss_100+0.681+0.563
deadline_miss_app (the product's headline form)+0.299+0.378

burst_max_ms vs loss_pct: +0.128 [+0.107, +0.150] adaptive, +0.149 [+0.127, +0.173] fixed. Holm p ≈ 0 on both. Confirmed.

And the mechanism shows itself in the difference between the two clients: on the fixed client, which cannot absorb delay, late_excess_mean (+0.629) rises above loss_pct (+0.540). Lateness magnitude matters most exactly where the client cannot buffer it away, which is what the mechanism predicts and is direct support for the original challenge.

E · What this changes

  1. The product is headlining the wrong figure. deadline_miss_app, a miss rate at the application's tick budget, is the worst deadline form in every analysis we have run. The longest hole in milliseconds beats it by +0.54 partial ρ on the adaptive client.
  2. The PRD's burst section was right and the metric bank ignored it. "1% arriving as a two-second hole is a lost gunfight" is the strongest measured claim in this whole programme, and burst structure was absent from the Stage 1 and 1b predictor set. That was my error and it cost two stages.
  3. **Lateness is vindicated as a mechanism and demoted as a rate.** The honest product line is that harm scales with how long you were without current information, whether that gap came from a packet being late or from its never arriving. Hole duration measures precisely that; a miss percentage does not.
  4. Stage 3 stays gated, on the grounds that the winning metric has now changed twice and should settle before human subjects are paid for.

Harness, seeds and raw JSON live beside these documents in the repository. Every figure here comes from simulated impairment; no human subjects were involved and no perception claim is made. Back to Bullettime.