Research · published whether it suits us or not

Proving the central claim — the programme

Status after 26 August 2026: Stages 0, 1, 1b, 1c and 1d complete.

The programme is staged so the cheap stages can kill the hypothesis before the expensive ones are paid for. That is not a formality: Stage 1 cost an afternoon and has just saved the cost of a 48-subject panel that would have been testing a claim its own screening does not support.


The claim, restated precisely

The PRD currently says no study compares late-packet rate against latency, jitter or conventional loss on the same traces. After the literature pass that is too strong in one direction and too weak in another, and should be replaced with:

The mechanism is standards consensus. ITU-T G.107 has folded late arrivals into effective loss since at least 2015 (Peff = Ppl + (1−Ppl)·Ppd); RFC 3611, 7002 and 8451 define discard metrics distinct from network loss; G.1072 and P.1203/P.1204 treat a frame past its deadline as lost; WebRTC exposes packetsDiscarded. What is unestablished is the comparative claim: no published work runs the head-to-head that would show a deadline-miss rate predicts user-reported problems better than the conventional metrics computed from the same traces.

That is the honest gap, it is narrower than the PRD's version, and it is still a real one.


Stage 0 — literature and preregistration · done

Two independent research runs, one Claude-family (free lane) and one Perplexity (~$2.00), in reports/. 83 cited sources between them. Both preregistrations written before their runs.

Stage 1 — automated screening, no human subjects · done, and it failed

1,280 simulated sessions per stage, non-circular game oracle, preregistered analysis. Full findings in RESULTS.md. Headline: the candidate lost to owd_mean in Stage 1 and to loss_pct in Stage 1b on a fresh seed with the oracle defect fixed. One exploratory form (deadline_miss_100, effective loss at a fixed 100 ms) beat conventional loss by a small, significant margin, and the application-tuned deadline was the worst-performing deadline form.

Gate: not passed.

Stage 1c / 1d — done. See RESULTS.md §A-E.

The rework below was carried out and it changed the answer. What follows is superseded and kept so the sequence is legible; item 1 (a staleness-aware damage measure) is still outstanding.

Stage 1c — the required rework · superseded by RESULTS.md §C-D

Three things, all identified by Stage 1 rather than guessed at:

  1. A damage measure that charges for staleness. The adaptive client absorbs constant delay entirely and is therefore blind to the cost of rendering a quarter-second in the past, which is a real cost to a real player. Add a staleness term to the oracle's damage and re-run. Neither existing client is right and this is the defect between them.
  2. Confirm or drop the deadline_miss_100 result on a third seed, preregistered, as its own primary rather than as an exploratory note. It is currently a post-hoc observation and counts for nothing until it survives that.
  3. A sweep over the deadline parameter, 4–250 ms, to establish whether there is any deadline at which the metric beats conventional loss conditionally, and where. If the answer is "only near 100 ms, and only marginally", the product's differentiator is Peff under a new name and the positioning must say so.

Gate: Stage 1c must show a conditional advantage over loss_pct with a 95% CI excluding zero, on a fresh seed, on a staleness-aware oracle. If it does not, the programme stops and the product drops the comparative claim permanently.

Stage 2 — emulation validity · scaffolded, not run

Replay the same impairment profiles through Linux tc-netem in a container and check the measured trace statistics agree with the simulator's. Both research runs independently flag that realised netem jitter can run systematically below the configured value, and that netem's independent per-packet delay draw manufactures reordering that no real FIFO queue produces. Our simulator uses an AR(1) delay process specifically to avoid the second, and measures 2.4% reordering; Stage 2 is what turns that from a design intention into a measurement.

Docker is available on this machine, so this is a day of work, not a procurement exercise.

Gate: simulator and netem must agree on loss rate, PDV percentiles and run-length distribution within a stated tolerance, or Stage 1's internal validity is not established and everything above is a claim about a model.

Stage 3 — crowdsourced subjective, voice · gated shut

ITU-T P.808, 120 crowd workers × 32 conditions = 3,840 ratings, power > 0.95. Cheap per rating and the right instrument for a first perception signal. Do not commission until Stage 1c passes.

Stage 4 — lab subjective · gated shut

ITU-T P.800 for voice (48 subjects × 64 conditions, power 0.84) and P.809 for gaming (48 subjects × 48 conditions, power 0.81). Detecting a 0.5 MOS difference with multiple-comparison control needs 40+ subjects per experiment. Mixed-effects models with per-subject random intercepts; Vuong and Clarke tests for the non-nested comparison; 10-fold subject-disjoint cross-validation.

The one design rule that carries from Stage 1: the outcome must not be produced by a model that already contains a playout buffer. Scoring a deadline metric against G.107 or G.1072 output is tautological, because those models already define loss to include discards. Ground truth has to be human ratings or application-level artefacts measured directly.


What happens to the product in the meantime

The gate the PRD already sets is the right one and it now has teeth:

Harness, seeds and raw JSON live beside these documents in the repository. Every figure here comes from simulated impairment; no human subjects were involved and no perception claim is made. Back to Bullettime.