perf(player): p0-1b perf tests for fps, scrub latency, and media sync drift (#400)

## Summary

Second slice of `P0-1` from the player perf proposal: plugs the three steady-state scenarios — sustained playback FPS, scrub latency, and media-sync drift — into the perf gate that landed in #399. Adds the multi-video fixture they all share, wires three new shards into CI, and seeds one new baseline (`droppedFramesMax`).

## Why

#399 stood up the harness and proved it with a single load-time scenario. By itself that's enough to catch regressions in initial composition setup, but it can't catch the things players actually fail at in production:

- **FPS regressions** — a render-loop change that drops the ticker from 60 to 45 fps still loads fast.
- **Scrub latency regressions** — the inline-vs-isolated split (#397) is exactly the kind of code path where a refactor can silently push everyone back to the postMessage round trip.
- **Media drift** — runtime mirror logic (#396 in this stack) and per-frame scheduling tweaks can both cause video to slip out of sync with the composition clock without producing a single console error.

Each of these is a target metric in the proposal with a concrete budget. This PR turns those budgets into gated CI signals and produces continuous data for them on every player/core/runtime change.

## What changed

### Fixture — `packages/player/tests/perf/fixtures/10-video-grid/`

- `index.html`: 10-second composition, 1920×1080, 30 fps, with 10 simultaneously-decoding video tiles in a 5×2 grid plus a subtle GSAP scale "breath" on each tile (so the rAF/RVFC loops have real work to do without GSAP dominating the budget the decoder needs).
- `sample.mp4`: small (~190 KB) clip checked in so the fixture is hermetic — no external CDN dependency, identical bytes on every run.
- Same `data-composition-id="main"` host pattern as `gsap-heavy`, so the existing harness loader works without changes.

### `02-fps.ts` — sustained playback frame rate

- Loads `10-video-grid`, calls `player.play()`, samples `requestAnimationFrame` callbacks inside the iframe for 5 s.
- Crucial sequencing: install the rAF sampler **before** `play()`, wait for `__player.isPlaying() === true`, **then reset the sample buffer** — otherwise the postMessage round-trip ramp-up window drags the average down by 5–10 fps.
- FPS = `(samples − 1) / (lastTs − firstTs in s)`; uses rAF timestamps (the same ones the compositor saw) rather than wall-clock `setTimeout`, so we're measuring real frame production.
- Dropped-frame definition matches Chrome DevTools: gap > 1.5× (1000/60 ms) ≈ 25 ms = "missed at least one vsync."
- Aggregation across runs: `min(fps)` and `max(droppedFrames)` — worst case wins, since the proposal asserts a floor on fps and a ceiling on drops.
- Emits `playback_fps_min` (higher-is-better, baseline `fpsMin = 55`) and `playback_dropped_frames_max` (lower-is-better, baseline `droppedFramesMax = 3`).

### `04-scrub.ts` — scrub latency, inline + isolated

- Loads `10-video-grid`, pauses, then issues 10 seek calls in two batches: first the synchronous **inline** path (`<hyperframes-player>`'s default same-origin `_trySyncSeek`), then the **isolated** path (forced by replacing `_trySyncSeek` with `() => false`, which makes the player fall back to the postMessage `_sendControl("seek")` bridge that cross-origin embeds and pre-#397 builds use).
- Inline runs first so the isolated mode's monkey-patch can't bleed back into the inline samples.
- Detection: a rAF watcher inside the iframe polls `__player.getTime()` until it's within `MATCH_TOLERANCE_S = 0.05 s` of the requested target. Tolerance exists because the postMessage bridge converts seconds → frame number → seconds, and that round-trip can introduce sub-frame quantization drift even for targets on the canonical fps grid.
- Timing: `performance.timeOrigin + performance.now()` in both contexts. `timeOrigin` is consistent across same-process frames, so `t1 − t0` is a true wall-clock latency, not a host-only or iframe-only stopwatch.
- Targets alternate forward/backward (`1.0, 7.0, 2.0, 8.0, 3.0, 9.0, 4.0, 6.0, 5.0, 0.5`) so no two consecutive seeks land near each other — protects the rAF watcher from matching against a stale `getTime()` value before the seek command is processed.
- Aggregation: `percentile(95)` across the pooled per-seek latencies from every run. With 10 seeks × 2 modes × 3 runs we get 30 samples per mode per CI shard, enough for a stable p95.
- Emits `scrub_latency_p95_inline_ms` (lower-is-better, baseline `scrubLatencyP95InlineMs = 33`) and `scrub_latency_p95_isolated_ms` (lower-is-better, baseline `scrubLatencyP95IsolatedMs = 80`).

### `05-drift.ts` — media sync drift

- Loads `10-video-grid`, plays 6 s, instruments **every** `video[data-start]` element with `requestVideoFrameCallback`. Each callback records `(compositionTime, actualMediaTime)` plus a snapshot of the clip transform (`clipStart`, `clipMediaStart`, `clipPlaybackRate`).
- Drift = `|actualMediaTime − ((compTime − clipStart) × clipPlaybackRate + clipMediaStart)|` — the same transform the runtime applies in `packages/core/src/runtime/media.ts`, snapshotted once at sampler install so the per-frame work is just subtract + multiply + abs.
- Sustain window is 6 s (not the proposal's 10 s) because the fixture composition is exactly 10 s long and we want headroom before the end-of-timeline pause/clamp behavior. With 10 videos × ~25 fps × 6 s we still pool ~1500 samples per run — more than enough for a stable p95.
- Same "reset buffer after play confirmed" gotcha as `02-fps.ts`: frames captured during the postMessage round-trip would compare a non-zero `mediaTime` against `getTime() === 0` and inflate drift by hundreds of ms.
- Aggregation: `max()` and `percentile(95)` across the pooled per-frame drifts. The proposal's max-drift ceiling of 500 ms is intentional — the runtime hard-resyncs when `|currentTime − relTime| > 0.5 s`, so a regression past 500 ms means the corrective resync kicked in and the viewer saw a jump.
- Emits `media_drift_max_ms` (lower-is-better, baseline `driftMaxMs = 500`) and `media_drift_p95_ms` (lower-is-better, baseline `driftP95Ms = 100`).

### Wiring

- `packages/player/tests/perf/index.ts`: add `fps`, `scrub`, `drift` to `ScenarioId`, `DEFAULT_RUNS`, the default scenario list (`--scenarios` defaults to all four), and three new dispatch branches.
- `packages/player/tests/perf/perf-gate.ts`: add `droppedFramesMax: number` to `PerfBaseline`. Other baseline keys for these scenarios were already seeded in #399.
- `packages/player/tests/perf/baseline.json`: add `droppedFramesMax: 3`.
- `.github/workflows/player-perf.yml`: three new matrix shards (`fps` / `scrub` / `drift`) at `runs: 3`. Same `paths-filter` and same artifact-upload pattern as the `load` shard, so the summary job aggregates them automatically.

## Methodology highlights

These three patterns recur in all three scenarios and are worth noting because they're load-bearing for the numbers we report:

1. **Reset buffer after play-confirmed.** The `play()` API is async (postMessage), so any samples captured before `__player.isPlaying() === true` belong to ramp-up, not steady-state. Both `02-fps` and `05-drift` clear `__perfRafSamples` / `__perfDriftSamples` *after* the wait. Without this, fps drops 5–10 and drift inflates by hundreds of ms.
2. **Iframe-side timing.** All three scenarios time inside the iframe (`performance.timeOrigin + performance.now()` for scrub, rAF/RVFC timestamps for fps/drift) rather than host-side. The iframe is what the user sees; host-side timing would conflate Puppeteer's IPC overhead with real player latency.
3. **Stop sampling before pause.** Sampler is deactivated *before* `pause()` is issued, so the pause command's postMessage round-trip can't perturb the tail of the measurement window.

## Test plan

- [x] Local: `bun run player:perf` runs all four scenarios end-to-end on the 10-video-grid fixture.
- [x] Each scenario produces metrics matching its declared `baselineKey` so `perf-gate.ts` can find them.
- [x] Typecheck, lint, format pass on the new files.
- [x] Existing player unit tests untouched (no production code changes in this PR).
- [ ] First CI run will confirm the new shards complete inside the workflow timeout and that the summary job picks up their `metrics.json` artifacts.

## Stack

Step `P0-1b` of the player perf proposal. Builds on:

- `P0-1a` (#399): the harness, runner, gate, and CI workflow this PR plugs new scenarios into.

Followed by:

- `P0-1c` (#401): `06-parity` — live playback frame vs. synchronously-seeked reference frame, compared via SSIM, on the existing `gsap-heavy` fixture from #399.
This commit is contained in:
Vance Ingalls
2026-04-22 18:10:08 -07:00
committed by GitHub
parent 10d2725b54
commit 6f05fabbf8
9 changed files with 1061 additions and 5 deletions
+9
View File
@@ -42,6 +42,15 @@ jobs:
- shard: load
scenarios: load
runs: "5"
- shard: fps
scenarios: fps
runs: "3"
- shard: scrub
scenarios: scrub
runs: "3"
- shard: drift
scenarios: drift
runs: "3"
steps:
- uses: actions/checkout@v4
+1 -1
View File
@@ -1,7 +1,7 @@
{
"compLoadColdP95Ms": 2000,
"compLoadWarmP95Ms": 1000,
"fpsMin": 55,
"compositionTimeAdvancementRatioMin": 0.95,
"scrubLatencyP95IsolatedMs": 80,
"scrubLatencyP95InlineMs": 33,
"driftMaxMs": 500,
@@ -0,0 +1,126 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<title>perf fixture: 10-video-grid</title>
<style>
:root {
color-scheme: dark;
}
html,
body {
margin: 0;
padding: 0;
background: #050714;
color: #e6e6f0;
font-family:
system-ui,
-apple-system,
sans-serif;
overflow: hidden;
}
#root {
position: relative;
width: 1920px;
height: 1080px;
display: grid;
grid-template-columns: repeat(5, 1fr);
grid-template-rows: repeat(2, 1fr);
gap: 8px;
padding: 8px;
box-sizing: border-box;
}
.tile {
position: relative;
background: #111827;
border-radius: 12px;
overflow: hidden;
box-shadow: 0 0 0 1px rgba(255, 255, 255, 0.05);
will-change: transform;
}
.tile video {
position: absolute;
inset: 0;
width: 100%;
height: 100%;
object-fit: cover;
}
.tile .label {
position: absolute;
top: 8px;
left: 8px;
z-index: 2;
font:
600 14px/1 system-ui,
sans-serif;
color: #fff;
background: rgba(0, 0, 0, 0.6);
padding: 4px 8px;
border-radius: 6px;
pointer-events: none;
}
</style>
<script src="/vendor/gsap.min.js"></script>
<script data-hyperframes-runtime="1" src="/vendor/hyperframe.runtime.iife.js"></script>
</head>
<body>
<div
id="root"
data-composition-id="main"
data-width="1920"
data-height="1080"
data-duration="10"
data-fps="30"
></div>
<script>
(function () {
var TILE_COUNT = 10;
var DURATION_SEC = 10;
var root = document.getElementById("root");
var tiles = [];
for (var i = 0; i < TILE_COUNT; i++) {
var tile = document.createElement("div");
tile.className = "tile";
tile.id = "tile-" + i;
var label = document.createElement("div");
label.className = "label";
label.textContent = "video " + (i + 1);
tile.appendChild(label);
var video = document.createElement("video");
video.id = "video-" + i;
video.setAttribute("data-start", "0");
video.setAttribute("data-duration", String(DURATION_SEC));
video.setAttribute("data-track-index", String(i));
video.setAttribute("src", "sample.mp4");
video.setAttribute("preload", "auto");
video.setAttribute("playsinline", "");
video.muted = true;
tile.appendChild(video);
root.appendChild(tile);
tiles.push(tile);
}
// Lightweight parent timeline so the player has a non-empty composition
// to drive. Each tile gets a subtle scale "breath" over the full
// duration — enough to keep GSAP scrubbing real properties without
// dominating the rAF budget that the video decoder needs.
var tl = gsap.timeline({ paused: true });
for (var j = 0; j < tiles.length; j++) {
tl.fromTo(
tiles[j],
{ scale: 0.985 },
{ scale: 1, duration: DURATION_SEC, ease: "sine.inOut" },
0,
);
}
window.__timelines = window.__timelines || {};
window.__timelines["main"] = tl;
})();
</script>
</body>
</html>
+65 -3
View File
@@ -29,7 +29,10 @@ import { execFileSync } from "node:child_process";
import { existsSync, mkdirSync, writeFileSync } from "node:fs";
import { dirname, resolve } from "node:path";
import { fileURLToPath } from "node:url";
import { runFps } from "./scenarios/02-fps.ts";
import { runLoad } from "./scenarios/03-load.ts";
import { runScrub } from "./scenarios/04-scrub.ts";
import { runDrift } from "./scenarios/05-drift.ts";
import { reportAndGate, type GateMode, type GateResult, type Metric } from "./perf-gate.ts";
import { launchBrowser } from "./runner.ts";
import { startServer } from "./server.ts";
@@ -38,7 +41,42 @@ const HERE = dirname(fileURLToPath(import.meta.url));
const RESULTS_DIR = resolve(HERE, "results");
const RESULTS_FILE = resolve(RESULTS_DIR, "metrics.json");
type ScenarioId = "load";
type ScenarioId = "load" | "fps" | "scrub" | "drift";
/**
* Per-scenario default `runs` value when the caller didn't pass `--runs`.
*
* Why `load` gets 5 runs and the others get 3:
*
* - `load` reports a single p95 over `runs` measurements, so each `run` is
* one sample. p95 over n=3 is mostly noise (the 95th percentile of three
* numbers is just `max`), so we bump it to 5. We considered 10 — but cold
* load is the slowest scenario in the shard (~2s × 5 runs × 2 fixtures =
* ~20s with disk cache cleared), and going to 10 would push the load shard
* past 30s of pure-measurement wall time per CI invocation.
* - `fps` aggregates as `min(ratio)` over runs — 3 runs gives us a worst-
* of-three signal, which is what we want for a floor metric. Adding more
* runs would only make the ratio strictly smaller (more chances to catch
* a stall) and shift the threshold toward false positives from runner
* contention rather than real regressions.
* - `scrub` and `drift` *pool* their per-run samples (10 seeks/run for
* scrub, ~1500 RVFC frames/run for drift) and compute the percentile over
* the pooled set. Their effective sample count for the percentile is
* `runs × samples_per_run`, not `runs`, so 3 runs already gives 30+ scrub
* samples and 4500+ drift samples per shard — well above the n≈30 rule of
* thumb for a stable p95.
*
* TODO(player-perf): revisit `fps: 3` once we have ~2 weeks of CI baseline
* data — if `min(ratio)` shows >5% inter-run variance attributable to runner
* jitter (not real player regressions), bump to 5 and tighten the
* `compositionTimeAdvancementRatioMin` baseline accordingly.
*/
const DEFAULT_RUNS: Record<ScenarioId, number> = {
load: 5,
fps: 3,
scrub: 3,
drift: 3,
};
type ResultsFile = {
schemaVersion: 1;
@@ -88,7 +126,7 @@ function parseArgs(argv: string[]): ParsedArgs {
// `mode` is consumed (measure logs regressions but never fails; enforce
// exits non-zero on regression).
mode: (process.env.PLAYER_PERF_MODE as GateMode) === "enforce" ? "enforce" : "measure",
scenarios: ["load"],
scenarios: ["load", "fps", "scrub", "drift"],
runs: null,
fixture: null,
headful: false,
@@ -150,7 +188,31 @@ async function main(): Promise<void> {
const m = await runLoad({
browser,
origin: server.origin,
runs: args.runs ?? 5,
runs: args.runs ?? DEFAULT_RUNS.load,
fixture: args.fixture,
});
metrics.push(...m);
} else if (scenario === "fps") {
const m = await runFps({
browser,
origin: server.origin,
runs: args.runs ?? DEFAULT_RUNS.fps,
fixture: args.fixture,
});
metrics.push(...m);
} else if (scenario === "scrub") {
const m = await runScrub({
browser,
origin: server.origin,
runs: args.runs ?? DEFAULT_RUNS.scrub,
fixture: args.fixture,
});
metrics.push(...m);
} else if (scenario === "drift") {
const m = await runDrift({
browser,
origin: server.origin,
runs: args.runs ?? DEFAULT_RUNS.drift,
fixture: args.fixture,
});
metrics.push(...m);
+10 -1
View File
@@ -31,7 +31,16 @@ export type Metric = {
export type PerfBaseline = {
compLoadColdP95Ms: number;
compLoadWarmP95Ms: number;
fpsMin: number;
/**
* Floor on `(compositionTime advanced) / (wallClock elapsed)` over a sustained
* playback window — see packages/player/tests/perf/scenarios/02-fps.ts. A
* healthy player keeps up with its intended speed and reads ~1.0; values
* below 1.0 mean the composition clock fell behind real time, which is the
* actual user-visible jank we want to gate against. Refresh-rate independent
* by construction, so it does not saturate to display refresh on high-Hz
* runners the way the previous `fpsMin` did. Direction: higher-is-better.
*/
compositionTimeAdvancementRatioMin: number;
scrubLatencyP95IsolatedMs: number;
scrubLatencyP95InlineMs: number;
driftMaxMs: number;
@@ -0,0 +1,236 @@
/**
* Scenario 02: sustained playback against the composition clock.
*
* Loads the 10-video-grid fixture, calls `player.play()`, then samples
* `__player.getTime()` at fixed wall-clock intervals for ~5 seconds. The
* emitted metric is the ratio of composition-time advanced to wall-clock
* elapsed:
*
* composition_time_advancement_ratio = (getTime(end) - getTime(start)) / wallSeconds
*
* This reads ~1.0 when the runtime is keeping up with its intended playback
* speed and falls below 1.0 when the player stalls — a slow video decoder, a
* blocked main thread, a GC pause, anything that prevents the composition
* clock from advancing at real-time. The metric is independent of the host
* display refresh rate by construction: both numerator and denominator are
* wall-clock timestamps, neither is a frame count, so a 60Hz, 120Hz, or 240Hz
* runner sees the same value for a healthy player.
*
* Why we replaced the previous rAF-based FPS metric:
* The original implementation counted `requestAnimationFrame` ticks per
* wall-clock second and asserted `fps >= 55`. On a 120Hz CI runner that
* reads ~120 fps regardless of whether the composition is actually
* advancing, so the gate passed even when the player was silently stalling.
* See PR #400 review (jrusso1020 + miguel-heygen) for the full discussion;
* this implementation follows jrusso1020's "first choice" recommendation.
*
* Per the proposal:
* Test 1: Playback frame rate (player-perf-fps)
* Load 10-video composition → play 5s → measure how well the player kept
* up with the composition clock.
*
* Methodology details:
* - We install the wall-clock sampler before calling `play()` so the very
* first post-play tick is captured. We then wait for `__player.isPlaying()`
* to flip true (the parent→iframe `play` message is async via postMessage)
* and *reset* the sample buffer, so the measurement window only contains
* samples taken while the runtime was actively playing the timeline.
* - Sampling cadence is 100ms (10 samples/sec). That's fine-grained enough
* to spot a half-second stall but coarse enough that the sampler itself
* has negligible overhead. With a 5s window we collect ~50 samples; the
* ratio is computed from the first and last sample's `getTime()` values.
* - We use `setInterval` (not rAF) on purpose: rAF cadence is the metric we
* are trying to *avoid* depending on. `setInterval` is wall-clock-driven.
*
* Outputs one metric:
* - composition_time_advancement_ratio_min
* (higher-is-better, baseline key compositionTimeAdvancementRatioMin)
*
* Aggregation: `min(ratio)` across runs because the proposal asserts a floor
* — the worst run is the one that gates against regressions.
*/
import type { Browser, Frame, Page } from "puppeteer-core";
import { loadHostPage } from "../runner.ts";
import type { Metric } from "../perf-gate.ts";
export type FpsScenarioOpts = {
browser: Browser;
origin: string;
/** Number of measurement runs. */
runs: number;
/** If null, runs the default fixture (10-video-grid). */
fixture: string | null;
};
const DEFAULT_FIXTURE = "10-video-grid";
const PLAYBACK_DURATION_MS = 5_000;
const SAMPLE_INTERVAL_MS = 100;
const PLAY_CONFIRM_TIMEOUT_MS = 5_000;
const FRAME_LOOKUP_TIMEOUT_MS = 5_000;
declare global {
interface Window {
/** (wallClockMs, compositionTimeSec) pairs collected by the sampler. */
__perfPlaySamples?: Array<{ wall: number; comp: number }>;
/** setInterval handle used by the sampler; cleared at the end of the window. */
__perfPlaySamplerHandle?: number;
/** Hyperframes runtime player API exposed inside the composition iframe. */
__player?: {
play: () => void;
pause: () => void;
seek: (timeSeconds: number) => void;
getTime: () => number;
getDuration: () => number;
isPlaying: () => boolean;
};
}
}
type RunResult = {
ratio: number;
compElapsedSec: number;
wallElapsedSec: number;
samples: number;
};
/**
* Find the iframe Puppeteer Frame that hosts the fixture composition. The
* `<hyperframes-player>` shell wraps an iframe whose URL is derived from the
* player's `src` attribute, so we match by path substring rather than full URL.
*/
async function getFixtureFrame(page: Page, fixture: string): Promise<Frame> {
const expected = `/fixtures/${fixture}/`;
const deadline = Date.now() + FRAME_LOOKUP_TIMEOUT_MS;
while (Date.now() < deadline) {
const frame = page.frames().find((f) => f.url().includes(expected));
if (frame) return frame;
await new Promise((r) => setTimeout(r, 50));
}
throw new Error(`[scenario:fps] fixture frame not found for "${fixture}" within timeout`);
}
async function runOnce(
opts: FpsScenarioOpts,
fixture: string,
idx: number,
total: number,
): Promise<RunResult> {
const ctx = await opts.browser.createBrowserContext();
try {
const page = await ctx.newPage();
const { duration } = await loadHostPage(page, opts.origin, { fixture });
const frame = await getFixtureFrame(page, fixture);
// Install the wall-clock sampler in the iframe context. We use setInterval
// because rAF cadence is exactly the host-display-dependent signal we are
// trying NOT to depend on; setInterval is driven by the event loop and
// gives us samples at fixed wall-clock cadence regardless of refresh rate.
await frame.evaluate((sampleIntervalMs: number) => {
window.__perfPlaySamples = [];
window.__perfPlaySamplerHandle = window.setInterval(() => {
const comp = window.__player?.getTime?.();
if (typeof comp !== "number" || !Number.isFinite(comp)) return;
window.__perfPlaySamples!.push({
wall: performance.timeOrigin + performance.now(),
comp,
});
}, sampleIntervalMs);
}, SAMPLE_INTERVAL_MS);
// Issue play from the host page (parent of the iframe). The player's
// public `play()` posts a control message into the iframe.
await page.evaluate(() => {
const el = document.getElementById("player") as (HTMLElement & { play: () => void }) | null;
if (!el) throw new Error("[scenario:fps] player element missing on host page");
el.play();
});
// Wait for the runtime to actually transition to playing — this is the
// signal that the postMessage round trip + timeline.play() finished.
await frame.waitForFunction(() => window.__player?.isPlaying?.() === true, {
timeout: PLAY_CONFIRM_TIMEOUT_MS,
});
// Reset samples now that playback is confirmed running. Anything captured
// before this point belongs to the ramp-up window (composition clock at
// 0, wall clock advancing) and would skew the ratio toward 0.
await frame.evaluate(() => {
window.__perfPlaySamples = [];
});
// Sustain playback for the measurement window.
await new Promise((r) => setTimeout(r, PLAYBACK_DURATION_MS));
// Stop the sampler and harvest the samples before pausing the runtime,
// so the pause command can't perturb the tail of the sample window.
const samples = (await frame.evaluate(() => {
if (window.__perfPlaySamplerHandle !== undefined) {
clearInterval(window.__perfPlaySamplerHandle);
window.__perfPlaySamplerHandle = undefined;
}
return window.__perfPlaySamples ?? [];
})) as Array<{ wall: number; comp: number }>;
await page.evaluate(() => {
const el = document.getElementById("player") as (HTMLElement & { pause: () => void }) | null;
el?.pause();
});
if (samples.length < 2) {
throw new Error(
`[scenario:fps] run ${idx + 1}/${total}: only ${samples.length} composition-clock samples captured (composition duration ${duration}s)`,
);
}
const first = samples[0]!;
const last = samples[samples.length - 1]!;
const wallElapsedSec = (last.wall - first.wall) / 1000;
const compElapsedSec = last.comp - first.comp;
const ratio = wallElapsedSec > 0 ? compElapsedSec / wallElapsedSec : 0;
console.log(
`[scenario:fps] run[${idx + 1}/${total}] ratio=${ratio.toFixed(4)} compElapsed=${compElapsedSec.toFixed(3)}s wallElapsed=${wallElapsedSec.toFixed(3)}s samples=${samples.length}`,
);
await page.close();
return {
ratio,
compElapsedSec,
wallElapsedSec,
samples: samples.length,
};
} finally {
await ctx.close();
}
}
export async function runFps(opts: FpsScenarioOpts): Promise<Metric[]> {
const fixture = opts.fixture ?? DEFAULT_FIXTURE;
const runs = Math.max(1, opts.runs);
console.log(
`[scenario:fps] fixture=${fixture} runs=${runs} window=${PLAYBACK_DURATION_MS}ms sampleInterval=${SAMPLE_INTERVAL_MS}ms`,
);
const ratios: number[] = [];
for (let i = 0; i < runs; i++) {
const result = await runOnce(opts, fixture, i, runs);
ratios.push(result.ratio);
}
// Worst run wins: the proposal asserts a floor on this ratio, so a single
// bad run (slow decoder, GC pause, host contention) is the one that gates.
const ratioMin = Math.min(...ratios);
console.log(`[scenario:fps] aggregate min ratio=${ratioMin.toFixed(4)} runs=${runs}`);
return [
{
name: "composition_time_advancement_ratio_min",
baselineKey: "compositionTimeAdvancementRatioMin",
value: ratioMin,
unit: "ratio",
direction: "higher-is-better",
samples: ratios,
},
];
}
@@ -0,0 +1,307 @@
/**
* Scenario 04: scrub latency.
*
* Loads the 10-video-grid fixture, pauses the player, then issues 10 seek
* calls in sequence — first through the synchronous "inline" path, then
* through the postMessage-driven "isolated" path — and measures the wall-clock
* latency from each `seek()` call to the first paint where the iframe's
* timeline reports the new time.
*
* Per the proposal:
* Test 2: Scrub latency (player-perf-scrub)
* Load composition → seek to 10 positions in sequence → measure time
* from seek() call to state update callback
* Assert: p95 < 80ms (isolated), p95 < 33ms (inline, Phase 4+)
*
* Methodology details:
* - Both modes are measured in the same page load. Inline runs first so
* the isolated mode's monkey-patch (forcing `_trySyncSeek` to return
* false) doesn't bleed into the inline samples.
* - "Inline" mode is the default behavior of `<hyperframes-player>` when the
* iframe is same-origin and exposes `__player.seek()` synchronously.
* `seek()` lands the new frame in the same task as the input event.
* - "Isolated" mode is forced by replacing the player element's
* `_trySyncSeek` method with `() => false`, which sends the player
* element through the postMessage bridge — exactly what cross-origin
* embeds and Phase 1 (pre-sync) builds did.
* - Detection is via a `requestAnimationFrame` watcher inside the iframe
* that polls `__player.getTime()` until it is within `MATCH_TOLERANCE_S`
* of the requested target. We use a tolerance because the postMessage
* bridge converts seconds → frame number → seconds, which can introduce
* sub-frame quantization drift even for targets on the canonical fps grid.
* - Timing uses `performance.timeOrigin + performance.now()` in both the
* host and iframe contexts. `timeOrigin` is consistent across same-process
* frames, so the difference is a true wall-clock measurement of latency.
* - Seek targets alternate forward/backward across the 10s composition so
* no two consecutive seeks land near each other; this avoids the rAF
* watcher matching against a stale `getTime()` value before the seek
* command is processed.
*
* Outputs two metrics:
* - scrub_latency_p95_inline_ms (lower-is-better, baseline scrubLatencyP95InlineMs)
* - scrub_latency_p95_isolated_ms (lower-is-better, baseline scrubLatencyP95IsolatedMs)
*
* Aggregation: percentile(95) is computed across the pooled per-seek
* latencies from every run. With 10 seeks per mode per run × 3 runs we get
* 30 samples per mode per CI shard, which is enough for a stable p95.
*/
import type { Browser, Frame, Page } from "puppeteer-core";
import { loadHostPage, percentile } from "../runner.ts";
import type { Metric } from "../perf-gate.ts";
export type ScrubScenarioOpts = {
browser: Browser;
origin: string;
/** Number of measurement runs. */
runs: number;
/** If null, runs the default fixture (10-video-grid). */
fixture: string | null;
};
const DEFAULT_FIXTURE = "10-video-grid";
/** Targets are seconds within the composition (10s duration). */
const SEEK_TARGETS: readonly number[] = [1.0, 7.0, 2.0, 8.0, 3.0, 9.0, 4.0, 6.0, 5.0, 0.5];
/**
* Tolerance window the rAF watcher uses to decide that the iframe's reported
* `__player.getTime()` matches the requested seek target. 50ms = 1.5 frames at
* 30fps, which absorbs three sources of expected slippage:
*
* 1. **Frame quantization on the postMessage path.** `_sendControl("seek")`
* converts seconds → integer frame number → seconds inside the runtime,
* so e.g. a target of 1.0s on a 30fps composition lands at frame 30 →
* 1.000s exactly, but a target of 1.005s lands at frame 30 → still
* 1.000s, a 5ms quantization error baked into the API itself.
* 2. **Sub-frame intra-clip clock advance.** Even with the iframe paused,
* between the `seek()` call landing and the next rAF tick, the runtime
* may have already nudged time by a fraction of a frame as part of
* finalizing the seek; `getTime()` reports the post-finalize value.
* 3. **Variable host load + browser jitter on CI.** GitHub runners share
* cores, so a noisy neighbor can delay the rAF tick that would otherwise
* register the match by tens of ms. Picking a tolerance much tighter
* than this would gate against runner contention rather than player
* regressions.
*
* The metric this scenario asserts is *latency to user-visible match*, not
* *exact equality of the reported time*, so a 50ms acceptance window is the
* intended behavior — but if we ever want to tighten this (e.g. to assert
* sub-frame precision on the inline path now that PR #397 documented it),
* this is the knob to turn. Configurability is deliberately deferred until
* we have a concrete second use case; YAGNI.
*
* TODO(player-perf): revisit this constant after P0-1b lands and we have ~2
* weeks of CI baseline data — if the inline-mode samples consistently cluster
* well below 50ms, drop this to e.g. 16ms (1 frame @ 60fps) and split the
* tolerance per mode (tighter for inline, current for isolated).
*/
const MATCH_TOLERANCE_S = 0.05;
/** Per-seek timeout; isolated p95 in the proposal is 80ms, so 1s is huge headroom. */
const SEEK_TIMEOUT_MS = 1_000;
const PAUSE_CONFIRM_TIMEOUT_MS = 5_000;
const FRAME_LOOKUP_TIMEOUT_MS = 5_000;
declare global {
interface Window {
/** Promise resolved by the iframe rAF watcher with the wall-clock t1 of the matching paint. */
__perfScrubAwait?: Promise<number>;
__player?: {
play: () => void;
pause: () => void;
seek: (timeSeconds: number) => void;
getTime: () => number;
getDuration: () => number;
isPlaying: () => boolean;
};
}
}
type Mode = "inline" | "isolated";
type RunResult = {
inlineLatencies: number[];
isolatedLatencies: number[];
};
/**
* Find the iframe Puppeteer Frame that hosts the fixture composition. Same
* helper as 02-fps.ts; duplicated locally so each scenario file is
* self-contained.
*/
async function getFixtureFrame(page: Page, fixture: string): Promise<Frame> {
const expected = `/fixtures/${fixture}/`;
const deadline = Date.now() + FRAME_LOOKUP_TIMEOUT_MS;
while (Date.now() < deadline) {
const frame = page.frames().find((f) => f.url().includes(expected));
if (frame) return frame;
await new Promise((r) => setTimeout(r, 50));
}
throw new Error(`[scenario:scrub] fixture frame not found for "${fixture}" within timeout`);
}
/**
* Measure a single seek's latency.
*
* Sequence:
* 1. Install a rAF watcher in the iframe that resolves with the wall-clock
* timestamp of the first paint where `__player.getTime()` is within
* tolerance of `target`. Promise is stashed on `window.__perfScrubAwait`.
* 2. Capture host wall-clock t0 and call `el.seek(target)` in the same task.
* 3. Await the iframe's resolved Promise (returns t1).
* 4. Latency = t1 - t0 (ms).
*/
async function measureSingleSeek(page: Page, frame: Frame, target: number): Promise<number> {
await frame.evaluate(
(target: number, tolerance: number, timeoutMs: number) => {
window.__perfScrubAwait = new Promise<number>((resolve, reject) => {
const deadlineWall = performance.timeOrigin + performance.now() + timeoutMs;
const tick = () => {
const wall = performance.timeOrigin + performance.now();
const time = window.__player?.getTime?.() ?? Number.NaN;
if (Number.isFinite(time) && Math.abs(time - target) < tolerance) {
resolve(wall);
return;
}
if (wall > deadlineWall) {
reject(new Error(`[scrub] timeout target=${target} last=${time}`));
return;
}
requestAnimationFrame(tick);
};
requestAnimationFrame(tick);
});
},
target,
MATCH_TOLERANCE_S,
SEEK_TIMEOUT_MS,
);
const t0Wall = await page.evaluate((targetSeconds: number) => {
const el = document.getElementById("player") as
| (HTMLElement & { seek: (t: number) => void })
| null;
if (!el) throw new Error("[scenario:scrub] player element missing on host page");
const wall = performance.timeOrigin + performance.now();
el.seek(targetSeconds);
return wall;
}, target);
// Puppeteer awaits the Promise we stashed on window and returns its resolved value.
const t1Wall = (await frame.evaluate(() => window.__perfScrubAwait as Promise<number>)) as number;
return t1Wall - t0Wall;
}
async function runScrubBatch(
page: Page,
frame: Frame,
mode: Mode,
idx: number,
total: number,
): Promise<number[]> {
const latencies: number[] = [];
for (const target of SEEK_TARGETS) {
const latency = await measureSingleSeek(page, frame, target);
latencies.push(latency);
}
const p95 = percentile(latencies, 95);
console.log(
`[scenario:scrub] run[${idx + 1}/${total}] mode=${mode} p95=${p95.toFixed(2)}ms n=${latencies.length}`,
);
return latencies;
}
async function runOnce(
opts: ScrubScenarioOpts,
fixture: string,
idx: number,
total: number,
): Promise<RunResult> {
const ctx = await opts.browser.createBrowserContext();
try {
const page = await ctx.newPage();
const { duration } = await loadHostPage(page, opts.origin, { fixture });
const requiredDuration = Math.max(...SEEK_TARGETS);
if (duration < requiredDuration) {
throw new Error(
`[scenario:scrub] fixture composition is ${duration.toFixed(2)}s but scrub targets require >= ${requiredDuration}s`,
);
}
const frame = await getFixtureFrame(page, fixture);
// Defensively pause: the host shell doesn't autoplay, but `pause()` also
// cancels any pending autoplay-on-ready behavior and guarantees the
// timeline isn't ticking under our seek measurements.
await page.evaluate(() => {
const el = document.getElementById("player") as (HTMLElement & { pause?: () => void }) | null;
el?.pause?.();
});
await frame.waitForFunction(() => window.__player?.isPlaying?.() === false, {
timeout: PAUSE_CONFIRM_TIMEOUT_MS,
});
// Inline mode first — the player's default `_trySyncSeek` path lands the
// seek synchronously when the iframe is same-origin (which it is here).
const inlineLatencies = await runScrubBatch(page, frame, "inline", idx, total);
// Force isolated mode by shadowing `_trySyncSeek` on the instance with
// a function that always reports failure. The fallback in `seek()` then
// sends the seek through `_sendControl("seek", { frame })`, which is the
// same path a cross-origin embed (or a Phase 1 build without sync seek)
// would take.
await page.evaluate(() => {
const el = document.getElementById("player") as
| (HTMLElement & { _trySyncSeek?: (t: number) => boolean })
| null;
if (!el) throw new Error("[scenario:scrub] player element missing on host page");
el._trySyncSeek = () => false;
});
const isolatedLatencies = await runScrubBatch(page, frame, "isolated", idx, total);
await page.close();
return { inlineLatencies, isolatedLatencies };
} finally {
await ctx.close();
}
}
export async function runScrub(opts: ScrubScenarioOpts): Promise<Metric[]> {
const fixture = opts.fixture ?? DEFAULT_FIXTURE;
const runs = Math.max(1, opts.runs);
console.log(
`[scenario:scrub] fixture=${fixture} runs=${runs} seeks_per_mode=${SEEK_TARGETS.length} tolerance=${(MATCH_TOLERANCE_S * 1000).toFixed(0)}ms`,
);
const allInline: number[] = [];
const allIsolated: number[] = [];
for (let i = 0; i < runs; i++) {
const result = await runOnce(opts, fixture, i, runs);
allInline.push(...result.inlineLatencies);
allIsolated.push(...result.isolatedLatencies);
}
const inlineP95 = percentile(allInline, 95);
const isolatedP95 = percentile(allIsolated, 95);
console.log(
`[scenario:scrub] aggregate inline_p95=${inlineP95.toFixed(2)}ms isolated_p95=${isolatedP95.toFixed(2)}ms (runs=${runs} samples_per_mode=${allInline.length})`,
);
return [
{
name: "scrub_latency_p95_inline_ms",
baselineKey: "scrubLatencyP95InlineMs",
value: inlineP95,
unit: "ms",
direction: "lower-is-better",
samples: allInline,
},
{
name: "scrub_latency_p95_isolated_ms",
baselineKey: "scrubLatencyP95IsolatedMs",
value: isolatedP95,
unit: "ms",
direction: "lower-is-better",
samples: allIsolated,
},
];
}
@@ -0,0 +1,307 @@
/**
* Scenario 05: media sync drift.
*
* Loads the 10-video-grid fixture, starts playback, and uses
* `requestVideoFrameCallback` on every video element to record
* (compositionTime, actualMediaTime) pairs for each decoded frame. Drift is
* the absolute difference between the *expected* media time (derived from the
* composition time using the runtime's clip transform) and the actual media
* time the decoder presented to the compositor.
*
* Per the proposal:
* Test 4: Media sync drift (player-perf-drift)
* Load 5-video composition → play for 10 seconds → on each RVFC callback,
* record drift between expected and actual media time
* Assert: max drift < 500ms, p95 drift < 100ms
*
* Methodology details:
* - We instrument *every* `video[data-start]` element in the fixture. The
* proposal called for 5 videos; the 10-video-grid gives us 10 streams in
* the same composition, which is a more conservative regression signal.
* - The expected media time uses the same transform the runtime applies in
* packages/core/src/runtime/media.ts:
*
* expectedMediaTime = (compositionTime - clip.start) * clip.playbackRate
* + clip.mediaStart
*
* We snapshot `clip.start` / `clip.mediaStart` / `clip.playbackRate` from
* each element's dataset + `defaultPlaybackRate` once when the sampler is
* installed, so the per-frame work is just a subtract + multiply + abs.
* - The runtime's media sync runs on a 50ms `setInterval`. Between syncs the
* video element's clock free-runs. The drift we measure here is the
* residual after that 50ms loop catches up — i.e. the user-visible glitch
* budget. The runtime hard-resyncs when |currentTime - relTime| > 0.5s
* (see media.ts), which is exactly the proposal's max-drift ceiling: a
* regression past 500ms means the corrective resync kicked in and the
* viewer saw a jump.
* - We install RVFC *before* calling play(), then reset the sample buffer
* once `__player.isPlaying()` flips true. Frames captured during the
* postMessage round-trip would compare a non-zero mediaTime against
* `getTime() === 0` and inflate drift to several hundred ms — same gotcha
* as 02-fps.ts.
* - Sustain window is 6s instead of the proposal's 10s because the fixture
* composition is exactly 10s long, and we want headroom before the
* end-of-timeline pause/clamp behavior. With 10 videos × ~25fps × 6s we
* still pool ~1500 samples per run, more than enough for a stable p95.
*
* Outputs two metrics:
* - media_drift_max_ms (lower-is-better, baseline driftMaxMs)
* - media_drift_p95_ms (lower-is-better, baseline driftP95Ms)
*
* Aggregation: max() and percentile(95) across the pooled per-frame drifts
* from every video in every run.
*/
import type { Browser, Frame, Page } from "puppeteer-core";
import { loadHostPage, percentile } from "../runner.ts";
import type { Metric } from "../perf-gate.ts";
export type DriftScenarioOpts = {
browser: Browser;
origin: string;
/** Number of measurement runs. */
runs: number;
/** If null, runs the default fixture (10-video-grid). */
fixture: string | null;
};
const DEFAULT_FIXTURE = "10-video-grid";
const PLAYBACK_DURATION_MS = 6_000;
const PLAY_CONFIRM_TIMEOUT_MS = 5_000;
const FRAME_LOOKUP_TIMEOUT_MS = 5_000;
type DriftSample = {
compTime: number;
actualMediaTime: number;
clipStart: number;
clipMediaStart: number;
clipPlaybackRate: number;
};
declare global {
interface Window {
/** RVFC samples collected by the iframe-side observer. */
__perfDriftSamples?: DriftSample[];
/** Set to false to stop sampling at the end of the measurement window. */
__perfDriftActive?: boolean;
__player?: {
play: () => void;
pause: () => void;
seek: (timeSeconds: number) => void;
getTime: () => number;
getDuration: () => number;
isPlaying: () => boolean;
};
}
}
type RunResult = {
drifts: number[];
videoCount: number;
};
/**
* Find the iframe Puppeteer Frame that hosts the fixture composition. Same
* helper as the other scenarios; duplicated locally so each scenario file is
* self-contained.
*/
async function getFixtureFrame(page: Page, fixture: string): Promise<Frame> {
const expected = `/fixtures/${fixture}/`;
const deadline = Date.now() + FRAME_LOOKUP_TIMEOUT_MS;
while (Date.now() < deadline) {
const frame = page.frames().find((f) => f.url().includes(expected));
if (frame) return frame;
await new Promise((r) => setTimeout(r, 50));
}
throw new Error(`[scenario:drift] fixture frame not found for "${fixture}" within timeout`);
}
async function runOnce(
opts: DriftScenarioOpts,
fixture: string,
idx: number,
total: number,
): Promise<RunResult> {
const ctx = await opts.browser.createBrowserContext();
try {
const page = await ctx.newPage();
const { duration } = await loadHostPage(page, opts.origin, { fixture });
const requiredDurationSec = PLAYBACK_DURATION_MS / 1000;
if (duration < requiredDurationSec) {
throw new Error(
`[scenario:drift] fixture composition is ${duration.toFixed(2)}s but drift sample window needs >= ${requiredDurationSec.toFixed(0)}s`,
);
}
const frame = await getFixtureFrame(page, fixture);
// Install RVFC on every `video[data-start]` element in the iframe. Each
// callback records the wall-clock-aligned (compositionTime, mediaTime)
// pair plus a snapshot of the clip transform so we can compute drift in
// node without re-querying the dataset on every frame.
const videoCount = (await frame.evaluate(() => {
window.__perfDriftSamples = [];
window.__perfDriftActive = true;
const videos = Array.from(document.querySelectorAll<HTMLVideoElement>("video[data-start]"));
type RvfcMetadata = { mediaTime: number; presentationTime: number };
type RvfcVideo = HTMLVideoElement & {
requestVideoFrameCallback?: (
cb: (now: DOMHighResTimeStamp, metadata: RvfcMetadata) => void,
) => number;
};
let installed = 0;
for (const video of videos) {
const rvfcVideo = video as RvfcVideo;
const rvfc = rvfcVideo.requestVideoFrameCallback;
// Headless Chrome supports RVFC; bail quietly on browsers that don't.
if (!rvfc) continue;
const clipStart = Number.parseFloat(video.dataset.start ?? "0") || 0;
const clipMediaStart =
Number.parseFloat(video.dataset.playbackStart ?? video.dataset.mediaStart ?? "0") || 0;
const rawRate = video.defaultPlaybackRate;
const clipPlaybackRate =
Number.isFinite(rawRate) && rawRate > 0 ? Math.max(0.1, Math.min(5, rawRate)) : 1;
const tick = (_now: DOMHighResTimeStamp, metadata: RvfcMetadata) => {
if (!window.__perfDriftActive) return;
const compTime = window.__player?.getTime?.() ?? Number.NaN;
if (Number.isFinite(compTime)) {
window.__perfDriftSamples!.push({
compTime,
actualMediaTime: metadata.mediaTime,
clipStart,
clipMediaStart,
clipPlaybackRate,
});
}
rvfc.call(video, tick);
};
rvfc.call(video, tick);
installed++;
}
return installed;
})) as number;
if (videoCount === 0) {
throw new Error(`[scenario:drift] fixture ${fixture} contains no video[data-start] elements`);
}
// Issue play from the host page; the player posts a control message into
// the iframe and the runtime starts the 50ms media sync poll.
await page.evaluate(() => {
const el = document.getElementById("player") as (HTMLElement & { play: () => void }) | null;
if (!el) throw new Error("[scenario:drift] player element missing on host page");
el.play();
});
// Wait for the runtime to confirm playing before we trust the samples.
await frame.waitForFunction(() => window.__player?.isPlaying?.() === true, {
timeout: PLAY_CONFIRM_TIMEOUT_MS,
});
// Reset the buffer now that playback is live. Anything captured during
// the postMessage round-trip would compare a non-zero mediaTime against
// `getTime() === 0` and bias drift up by hundreds of ms.
await frame.evaluate(() => {
window.__perfDriftSamples = [];
});
await new Promise((r) => setTimeout(r, PLAYBACK_DURATION_MS));
// Stop sampling first, then pause. Same ordering as 02-fps.ts so the
// pause command can't perturb the tail of the measurement window.
const samples = (await frame.evaluate(() => {
window.__perfDriftActive = false;
return window.__perfDriftSamples ?? [];
})) as DriftSample[];
await page.evaluate(() => {
const el = document.getElementById("player") as (HTMLElement & { pause: () => void }) | null;
el?.pause();
});
if (samples.length === 0) {
throw new Error(
`[scenario:drift] run ${idx + 1}/${total}: zero RVFC samples captured (videos=${videoCount}, duration=${duration.toFixed(2)}s)`,
);
}
// Apply the runtime's transform to derive the expected media time, then
// compare against the actual media time the decoder presented. Convert
// to ms here so the gate threshold (driftMaxMs / driftP95Ms) compares
// apples-to-apples.
const drifts: number[] = [];
for (const s of samples) {
const expectedMediaTime = (s.compTime - s.clipStart) * s.clipPlaybackRate + s.clipMediaStart;
const driftMs = Math.abs(s.actualMediaTime - expectedMediaTime) * 1000;
drifts.push(driftMs);
}
const max = Math.max(...drifts);
const p95 = percentile(drifts, 95);
console.log(
`[scenario:drift] run[${idx + 1}/${total}] max=${max.toFixed(2)}ms p95=${p95.toFixed(2)}ms videos=${videoCount} samples=${samples.length}`,
);
await page.close();
return { drifts, videoCount };
} finally {
await ctx.close();
}
}
export async function runDrift(opts: DriftScenarioOpts): Promise<Metric[]> {
const fixture = opts.fixture ?? DEFAULT_FIXTURE;
const runs = Math.max(1, opts.runs);
console.log(`[scenario:drift] fixture=${fixture} runs=${runs} window=${PLAYBACK_DURATION_MS}ms`);
const allDrifts: number[] = [];
let lastVideoCount = 0;
for (let i = 0; i < runs; i++) {
const result = await runOnce(opts, fixture, i, runs);
allDrifts.push(...result.drifts);
lastVideoCount = result.videoCount;
}
// Worst case wins for max; p95 is computed across the pooled per-frame
// drifts from every video in every run. The proposal asserts max < 500ms
// and p95 < 100ms, so a single bad sample legitimately gates the build.
const maxDrift = Math.max(...allDrifts);
const p95Drift = percentile(allDrifts, 95);
// Coefficient of variation (stddev / mean) is logged here as a soft signal
// we can eyeball in CI output. We deliberately do NOT gate on it — the
// baseline asserts absolute thresholds (max, p95), and the underlying
// distribution is heavy-tailed (most frames are sub-50ms, occasional ones
// spike during the 50ms media-sync interval). But CV is a useful early
// warning: if it climbs significantly across CI runs while max + p95 stay
// green, our jitter assumptions about the runtime's resync loop have
// shifted (e.g. if media.ts changes its 50ms `setInterval` cadence) and
// we should revisit the baselines before they start producing flakes.
// TODO(player-perf): once we have ~2 weeks of CI baseline data, decide
// whether to publish CV as a tracked-but-ungated metric in baseline.json
// alongside max + p95, or wire it into the Slack regression report.
const meanDrift = allDrifts.reduce((a, b) => a + b, 0) / allDrifts.length;
const variance = allDrifts.reduce((acc, d) => acc + (d - meanDrift) ** 2, 0) / allDrifts.length;
const stddev = Math.sqrt(variance);
const cv = meanDrift > 0 ? stddev / meanDrift : 0;
console.log(
`[scenario:drift] aggregate max=${maxDrift.toFixed(2)}ms p95=${p95Drift.toFixed(2)}ms mean=${meanDrift.toFixed(2)}ms cv=${cv.toFixed(3)} videos=${lastVideoCount} samples=${allDrifts.length} runs=${runs}`,
);
return [
{
name: "media_drift_max_ms",
baselineKey: "driftMaxMs",
value: maxDrift,
unit: "ms",
direction: "lower-is-better",
samples: allDrifts,
},
{
name: "media_drift_p95_ms",
baselineKey: "driftP95Ms",
value: p95Drift,
unit: "ms",
direction: "lower-is-better",
samples: allDrifts,
},
];
}