Proving the central claim — the programme
Status after 26 August 2026: Stages 0, 1, 1b, 1c and 1d complete.
- The mechanism is confirmed and near-deterministic: staleness against client error, Spearman +0.950 at the event level, graded over three orders of magnitude.
- The miss-rate form the product headlines failed three times and is the worst-performing deadline form we have measured.
- The metric that wins, confirmed on a fresh seed on both client models, is the **longest continuous hole in milliseconds** (
burst_max_ms), beating conventional loss by +0.13 to +0.15 partial rho, Holm p ~ 0. - Stages 3 and 4 remain gated, now on the grounds that the winning metric changed twice and should settle before subjects are paid for.
The programme is staged so the cheap stages can kill the hypothesis before the expensive ones are paid for. That is not a formality: Stage 1 cost an afternoon and has just saved the cost of a 48-subject panel that would have been testing a claim its own screening does not support.
The claim, restated precisely
The PRD currently says no study compares late-packet rate against latency, jitter or conventional loss on the same traces. After the literature pass that is too strong in one direction and too weak in another, and should be replaced with:
The mechanism is standards consensus. ITU-T G.107 has folded late arrivals into effective loss since at least 2015 (Peff = Ppl + (1−Ppl)·Ppd); RFC 3611, 7002 and 8451 define discard metrics distinct from network loss; G.1072 and P.1203/P.1204 treat a frame past its deadline as lost; WebRTC exposespacketsDiscarded. What is unestablished is the comparative claim: no published work runs the head-to-head that would show a deadline-miss rate predicts user-reported problems better than the conventional metrics computed from the same traces.
That is the honest gap, it is narrower than the PRD's version, and it is still a real one.
Stage 0 — literature and preregistration · done
Two independent research runs, one Claude-family (free lane) and one Perplexity (~$2.00), in reports/. 83 cited sources between them. Both preregistrations written before their runs.
Stage 1 — automated screening, no human subjects · done, and it failed
1,280 simulated sessions per stage, non-circular game oracle, preregistered analysis. Full findings in RESULTS.md. Headline: the candidate lost to owd_mean in Stage 1 and to loss_pct in Stage 1b on a fresh seed with the oracle defect fixed. One exploratory form (deadline_miss_100, effective loss at a fixed 100 ms) beat conventional loss by a small, significant margin, and the application-tuned deadline was the worst-performing deadline form.
Gate: not passed.
Stage 1c / 1d — done. See RESULTS.md §A-E.
The rework below was carried out and it changed the answer. What follows is superseded and kept so the sequence is legible; item 1 (a staleness-aware damage measure) is still outstanding.
Stage 1c — the required rework · superseded by RESULTS.md §C-D
Three things, all identified by Stage 1 rather than guessed at:
- A damage measure that charges for staleness. The adaptive client absorbs constant delay entirely and is therefore blind to the cost of rendering a quarter-second in the past, which is a real cost to a real player. Add a staleness term to the oracle's damage and re-run. Neither existing client is right and this is the defect between them.
- Confirm or drop the
deadline_miss_100result on a third seed, preregistered, as its own primary rather than as an exploratory note. It is currently a post-hoc observation and counts for nothing until it survives that. - A sweep over the deadline parameter, 4–250 ms, to establish whether there is any deadline at which the metric beats conventional loss conditionally, and where. If the answer is "only near 100 ms, and only marginally", the product's differentiator is
Peffunder a new name and the positioning must say so.
Gate: Stage 1c must show a conditional advantage over loss_pct with a 95% CI excluding zero, on a fresh seed, on a staleness-aware oracle. If it does not, the programme stops and the product drops the comparative claim permanently.
Stage 2 — emulation validity · scaffolded, not run
Replay the same impairment profiles through Linux tc-netem in a container and check the measured trace statistics agree with the simulator's. Both research runs independently flag that realised netem jitter can run systematically below the configured value, and that netem's independent per-packet delay draw manufactures reordering that no real FIFO queue produces. Our simulator uses an AR(1) delay process specifically to avoid the second, and measures 2.4% reordering; Stage 2 is what turns that from a design intention into a measurement.
Docker is available on this machine, so this is a day of work, not a procurement exercise.
Gate: simulator and netem must agree on loss rate, PDV percentiles and run-length distribution within a stated tolerance, or Stage 1's internal validity is not established and everything above is a claim about a model.
Stage 3 — crowdsourced subjective, voice · gated shut
ITU-T P.808, 120 crowd workers × 32 conditions = 3,840 ratings, power > 0.95. Cheap per rating and the right instrument for a first perception signal. Do not commission until Stage 1c passes.
Stage 4 — lab subjective · gated shut
ITU-T P.800 for voice (48 subjects × 64 conditions, power 0.84) and P.809 for gaming (48 subjects × 48 conditions, power 0.81). Detecting a 0.5 MOS difference with multiple-comparison control needs 40+ subjects per experiment. Mixed-effects models with per-subject random intercepts; Vuong and Clarke tests for the non-nested comparison; 10-fold subject-disjoint cross-validation.
The one design rule that carries from Stage 1: the outcome must not be produced by a model that already contains a playout buffer. Scoring a deadline metric against G.107 or G.1072 output is tautological, because those models already define loss to include discards. Ground truth has to be human ratings or application-level artefacts measured directly.
What happens to the product in the meantime
The gate the PRD already sets is the right one and it now has teeth:
- The mechanism may be described, and cited to the standards, because that is what it is.
- The comparative claim may not be made in marketing, in the app, or in an export, until Stage 1c and Stage 3 pass. The marketing page already carries an open-question section; it now needs to say that the claim was tested and did not hold, not merely that it is untested.
- The per-application deadline is, on current evidence, the weakest form of the metric rather than the strongest. It should stop being the headline until Stage 1c's parameter sweep says otherwise.
- The survival curve stays, as a presentation. "We show you the whole curve instead of guessing your buffer depth" is a claim about honesty in presentation and does not depend on the curve's integral being a good predictor, which it is not.
Harness, seeds and raw JSON live beside these documents in the repository. Every figure here comes from simulated impairment; no human subjects were involved and no perception claim is made. Back to Bullettime.