mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-04 16:42:27 +00:00
## Summary Second slice of `P0-1` from the player perf proposal: plugs the three steady-state scenarios — sustained playback FPS, scrub latency, and media-sync drift — into the perf gate that landed in #399. Adds the multi-video fixture they all share, wires three new shards into CI, and seeds one new baseline (`droppedFramesMax`). ## Why #399 stood up the harness and proved it with a single load-time scenario. By itself that's enough to catch regressions in initial composition setup, but it can't catch the things players actually fail at in production: - **FPS regressions** — a render-loop change that drops the ticker from 60 to 45 fps still loads fast. - **Scrub latency regressions** — the inline-vs-isolated split (#397) is exactly the kind of code path where a refactor can silently push everyone back to the postMessage round trip. - **Media drift** — runtime mirror logic (#396 in this stack) and per-frame scheduling tweaks can both cause video to slip out of sync with the composition clock without producing a single console error. Each of these is a target metric in the proposal with a concrete budget. This PR turns those budgets into gated CI signals and produces continuous data for them on every player/core/runtime change. ## What changed ### Fixture — `packages/player/tests/perf/fixtures/10-video-grid/` - `index.html`: 10-second composition, 1920×1080, 30 fps, with 10 simultaneously-decoding video tiles in a 5×2 grid plus a subtle GSAP scale "breath" on each tile (so the rAF/RVFC loops have real work to do without GSAP dominating the budget the decoder needs). - `sample.mp4`: small (~190 KB) clip checked in so the fixture is hermetic — no external CDN dependency, identical bytes on every run. - Same `data-composition-id="main"` host pattern as `gsap-heavy`, so the existing harness loader works without changes. ### `02-fps.ts` — sustained playback frame rate - Loads `10-video-grid`, calls `player.play()`, samples `requestAnimationFrame` callbacks inside the iframe for 5 s. - Crucial sequencing: install the rAF sampler **before** `play()`, wait for `__player.isPlaying() === true`, **then reset the sample buffer** — otherwise the postMessage round-trip ramp-up window drags the average down by 5–10 fps. - FPS = `(samples − 1) / (lastTs − firstTs in s)`; uses rAF timestamps (the same ones the compositor saw) rather than wall-clock `setTimeout`, so we're measuring real frame production. - Dropped-frame definition matches Chrome DevTools: gap > 1.5× (1000/60 ms) ≈ 25 ms = "missed at least one vsync." - Aggregation across runs: `min(fps)` and `max(droppedFrames)` — worst case wins, since the proposal asserts a floor on fps and a ceiling on drops. - Emits `playback_fps_min` (higher-is-better, baseline `fpsMin = 55`) and `playback_dropped_frames_max` (lower-is-better, baseline `droppedFramesMax = 3`). ### `04-scrub.ts` — scrub latency, inline + isolated - Loads `10-video-grid`, pauses, then issues 10 seek calls in two batches: first the synchronous **inline** path (`<hyperframes-player>`'s default same-origin `_trySyncSeek`), then the **isolated** path (forced by replacing `_trySyncSeek` with `() => false`, which makes the player fall back to the postMessage `_sendControl("seek")` bridge that cross-origin embeds and pre-#397 builds use). - Inline runs first so the isolated mode's monkey-patch can't bleed back into the inline samples. - Detection: a rAF watcher inside the iframe polls `__player.getTime()` until it's within `MATCH_TOLERANCE_S = 0.05 s` of the requested target. Tolerance exists because the postMessage bridge converts seconds → frame number → seconds, and that round-trip can introduce sub-frame quantization drift even for targets on the canonical fps grid. - Timing: `performance.timeOrigin + performance.now()` in both contexts. `timeOrigin` is consistent across same-process frames, so `t1 − t0` is a true wall-clock latency, not a host-only or iframe-only stopwatch. - Targets alternate forward/backward (`1.0, 7.0, 2.0, 8.0, 3.0, 9.0, 4.0, 6.0, 5.0, 0.5`) so no two consecutive seeks land near each other — protects the rAF watcher from matching against a stale `getTime()` value before the seek command is processed. - Aggregation: `percentile(95)` across the pooled per-seek latencies from every run. With 10 seeks × 2 modes × 3 runs we get 30 samples per mode per CI shard, enough for a stable p95. - Emits `scrub_latency_p95_inline_ms` (lower-is-better, baseline `scrubLatencyP95InlineMs = 33`) and `scrub_latency_p95_isolated_ms` (lower-is-better, baseline `scrubLatencyP95IsolatedMs = 80`). ### `05-drift.ts` — media sync drift - Loads `10-video-grid`, plays 6 s, instruments **every** `video[data-start]` element with `requestVideoFrameCallback`. Each callback records `(compositionTime, actualMediaTime)` plus a snapshot of the clip transform (`clipStart`, `clipMediaStart`, `clipPlaybackRate`). - Drift = `|actualMediaTime − ((compTime − clipStart) × clipPlaybackRate + clipMediaStart)|` — the same transform the runtime applies in `packages/core/src/runtime/media.ts`, snapshotted once at sampler install so the per-frame work is just subtract + multiply + abs. - Sustain window is 6 s (not the proposal's 10 s) because the fixture composition is exactly 10 s long and we want headroom before the end-of-timeline pause/clamp behavior. With 10 videos × ~25 fps × 6 s we still pool ~1500 samples per run — more than enough for a stable p95. - Same "reset buffer after play confirmed" gotcha as `02-fps.ts`: frames captured during the postMessage round-trip would compare a non-zero `mediaTime` against `getTime() === 0` and inflate drift by hundreds of ms. - Aggregation: `max()` and `percentile(95)` across the pooled per-frame drifts. The proposal's max-drift ceiling of 500 ms is intentional — the runtime hard-resyncs when `|currentTime − relTime| > 0.5 s`, so a regression past 500 ms means the corrective resync kicked in and the viewer saw a jump. - Emits `media_drift_max_ms` (lower-is-better, baseline `driftMaxMs = 500`) and `media_drift_p95_ms` (lower-is-better, baseline `driftP95Ms = 100`). ### Wiring - `packages/player/tests/perf/index.ts`: add `fps`, `scrub`, `drift` to `ScenarioId`, `DEFAULT_RUNS`, the default scenario list (`--scenarios` defaults to all four), and three new dispatch branches. - `packages/player/tests/perf/perf-gate.ts`: add `droppedFramesMax: number` to `PerfBaseline`. Other baseline keys for these scenarios were already seeded in #399. - `packages/player/tests/perf/baseline.json`: add `droppedFramesMax: 3`. - `.github/workflows/player-perf.yml`: three new matrix shards (`fps` / `scrub` / `drift`) at `runs: 3`. Same `paths-filter` and same artifact-upload pattern as the `load` shard, so the summary job aggregates them automatically. ## Methodology highlights These three patterns recur in all three scenarios and are worth noting because they're load-bearing for the numbers we report: 1. **Reset buffer after play-confirmed.** The `play()` API is async (postMessage), so any samples captured before `__player.isPlaying() === true` belong to ramp-up, not steady-state. Both `02-fps` and `05-drift` clear `__perfRafSamples` / `__perfDriftSamples` *after* the wait. Without this, fps drops 5–10 and drift inflates by hundreds of ms. 2. **Iframe-side timing.** All three scenarios time inside the iframe (`performance.timeOrigin + performance.now()` for scrub, rAF/RVFC timestamps for fps/drift) rather than host-side. The iframe is what the user sees; host-side timing would conflate Puppeteer's IPC overhead with real player latency. 3. **Stop sampling before pause.** Sampler is deactivated *before* `pause()` is issued, so the pause command's postMessage round-trip can't perturb the tail of the measurement window. ## Test plan - [x] Local: `bun run player:perf` runs all four scenarios end-to-end on the 10-video-grid fixture. - [x] Each scenario produces metrics matching its declared `baselineKey` so `perf-gate.ts` can find them. - [x] Typecheck, lint, format pass on the new files. - [x] Existing player unit tests untouched (no production code changes in this PR). - [ ] First CI run will confirm the new shards complete inside the workflow timeout and that the summary job picks up their `metrics.json` artifacts. ## Stack Step `P0-1b` of the player perf proposal. Builds on: - `P0-1a` (#399): the harness, runner, gate, and CI workflow this PR plugs new scenarios into. Followed by: - `P0-1c` (#401): `06-parity` — live playback frame vs. synchronously-seeked reference frame, compared via SSIM, on the existing `gsap-heavy` fixture from #399.
308 lines
12 KiB
TypeScript
308 lines
12 KiB
TypeScript
/**
|
||
* Scenario 04: scrub latency.
|
||
*
|
||
* Loads the 10-video-grid fixture, pauses the player, then issues 10 seek
|
||
* calls in sequence — first through the synchronous "inline" path, then
|
||
* through the postMessage-driven "isolated" path — and measures the wall-clock
|
||
* latency from each `seek()` call to the first paint where the iframe's
|
||
* timeline reports the new time.
|
||
*
|
||
* Per the proposal:
|
||
* Test 2: Scrub latency (player-perf-scrub)
|
||
* Load composition → seek to 10 positions in sequence → measure time
|
||
* from seek() call to state update callback
|
||
* Assert: p95 < 80ms (isolated), p95 < 33ms (inline, Phase 4+)
|
||
*
|
||
* Methodology details:
|
||
* - Both modes are measured in the same page load. Inline runs first so
|
||
* the isolated mode's monkey-patch (forcing `_trySyncSeek` to return
|
||
* false) doesn't bleed into the inline samples.
|
||
* - "Inline" mode is the default behavior of `<hyperframes-player>` when the
|
||
* iframe is same-origin and exposes `__player.seek()` synchronously.
|
||
* `seek()` lands the new frame in the same task as the input event.
|
||
* - "Isolated" mode is forced by replacing the player element's
|
||
* `_trySyncSeek` method with `() => false`, which sends the player
|
||
* element through the postMessage bridge — exactly what cross-origin
|
||
* embeds and Phase 1 (pre-sync) builds did.
|
||
* - Detection is via a `requestAnimationFrame` watcher inside the iframe
|
||
* that polls `__player.getTime()` until it is within `MATCH_TOLERANCE_S`
|
||
* of the requested target. We use a tolerance because the postMessage
|
||
* bridge converts seconds → frame number → seconds, which can introduce
|
||
* sub-frame quantization drift even for targets on the canonical fps grid.
|
||
* - Timing uses `performance.timeOrigin + performance.now()` in both the
|
||
* host and iframe contexts. `timeOrigin` is consistent across same-process
|
||
* frames, so the difference is a true wall-clock measurement of latency.
|
||
* - Seek targets alternate forward/backward across the 10s composition so
|
||
* no two consecutive seeks land near each other; this avoids the rAF
|
||
* watcher matching against a stale `getTime()` value before the seek
|
||
* command is processed.
|
||
*
|
||
* Outputs two metrics:
|
||
* - scrub_latency_p95_inline_ms (lower-is-better, baseline scrubLatencyP95InlineMs)
|
||
* - scrub_latency_p95_isolated_ms (lower-is-better, baseline scrubLatencyP95IsolatedMs)
|
||
*
|
||
* Aggregation: percentile(95) is computed across the pooled per-seek
|
||
* latencies from every run. With 10 seeks per mode per run × 3 runs we get
|
||
* 30 samples per mode per CI shard, which is enough for a stable p95.
|
||
*/
|
||
|
||
import type { Browser, Frame, Page } from "puppeteer-core";
|
||
import { loadHostPage, percentile } from "../runner.ts";
|
||
import type { Metric } from "../perf-gate.ts";
|
||
|
||
export type ScrubScenarioOpts = {
|
||
browser: Browser;
|
||
origin: string;
|
||
/** Number of measurement runs. */
|
||
runs: number;
|
||
/** If null, runs the default fixture (10-video-grid). */
|
||
fixture: string | null;
|
||
};
|
||
|
||
const DEFAULT_FIXTURE = "10-video-grid";
|
||
/** Targets are seconds within the composition (10s duration). */
|
||
const SEEK_TARGETS: readonly number[] = [1.0, 7.0, 2.0, 8.0, 3.0, 9.0, 4.0, 6.0, 5.0, 0.5];
|
||
/**
|
||
* Tolerance window the rAF watcher uses to decide that the iframe's reported
|
||
* `__player.getTime()` matches the requested seek target. 50ms = 1.5 frames at
|
||
* 30fps, which absorbs three sources of expected slippage:
|
||
*
|
||
* 1. **Frame quantization on the postMessage path.** `_sendControl("seek")`
|
||
* converts seconds → integer frame number → seconds inside the runtime,
|
||
* so e.g. a target of 1.0s on a 30fps composition lands at frame 30 →
|
||
* 1.000s exactly, but a target of 1.005s lands at frame 30 → still
|
||
* 1.000s, a 5ms quantization error baked into the API itself.
|
||
* 2. **Sub-frame intra-clip clock advance.** Even with the iframe paused,
|
||
* between the `seek()` call landing and the next rAF tick, the runtime
|
||
* may have already nudged time by a fraction of a frame as part of
|
||
* finalizing the seek; `getTime()` reports the post-finalize value.
|
||
* 3. **Variable host load + browser jitter on CI.** GitHub runners share
|
||
* cores, so a noisy neighbor can delay the rAF tick that would otherwise
|
||
* register the match by tens of ms. Picking a tolerance much tighter
|
||
* than this would gate against runner contention rather than player
|
||
* regressions.
|
||
*
|
||
* The metric this scenario asserts is *latency to user-visible match*, not
|
||
* *exact equality of the reported time*, so a 50ms acceptance window is the
|
||
* intended behavior — but if we ever want to tighten this (e.g. to assert
|
||
* sub-frame precision on the inline path now that PR #397 documented it),
|
||
* this is the knob to turn. Configurability is deliberately deferred until
|
||
* we have a concrete second use case; YAGNI.
|
||
*
|
||
* TODO(player-perf): revisit this constant after P0-1b lands and we have ~2
|
||
* weeks of CI baseline data — if the inline-mode samples consistently cluster
|
||
* well below 50ms, drop this to e.g. 16ms (1 frame @ 60fps) and split the
|
||
* tolerance per mode (tighter for inline, current for isolated).
|
||
*/
|
||
const MATCH_TOLERANCE_S = 0.05;
|
||
/** Per-seek timeout; isolated p95 in the proposal is 80ms, so 1s is huge headroom. */
|
||
const SEEK_TIMEOUT_MS = 1_000;
|
||
const PAUSE_CONFIRM_TIMEOUT_MS = 5_000;
|
||
const FRAME_LOOKUP_TIMEOUT_MS = 5_000;
|
||
|
||
declare global {
|
||
interface Window {
|
||
/** Promise resolved by the iframe rAF watcher with the wall-clock t1 of the matching paint. */
|
||
__perfScrubAwait?: Promise<number>;
|
||
__player?: {
|
||
play: () => void;
|
||
pause: () => void;
|
||
seek: (timeSeconds: number) => void;
|
||
getTime: () => number;
|
||
getDuration: () => number;
|
||
isPlaying: () => boolean;
|
||
};
|
||
}
|
||
}
|
||
|
||
type Mode = "inline" | "isolated";
|
||
|
||
type RunResult = {
|
||
inlineLatencies: number[];
|
||
isolatedLatencies: number[];
|
||
};
|
||
|
||
/**
|
||
* Find the iframe Puppeteer Frame that hosts the fixture composition. Same
|
||
* helper as 02-fps.ts; duplicated locally so each scenario file is
|
||
* self-contained.
|
||
*/
|
||
async function getFixtureFrame(page: Page, fixture: string): Promise<Frame> {
|
||
const expected = `/fixtures/${fixture}/`;
|
||
const deadline = Date.now() + FRAME_LOOKUP_TIMEOUT_MS;
|
||
while (Date.now() < deadline) {
|
||
const frame = page.frames().find((f) => f.url().includes(expected));
|
||
if (frame) return frame;
|
||
await new Promise((r) => setTimeout(r, 50));
|
||
}
|
||
throw new Error(`[scenario:scrub] fixture frame not found for "${fixture}" within timeout`);
|
||
}
|
||
|
||
/**
|
||
* Measure a single seek's latency.
|
||
*
|
||
* Sequence:
|
||
* 1. Install a rAF watcher in the iframe that resolves with the wall-clock
|
||
* timestamp of the first paint where `__player.getTime()` is within
|
||
* tolerance of `target`. Promise is stashed on `window.__perfScrubAwait`.
|
||
* 2. Capture host wall-clock t0 and call `el.seek(target)` in the same task.
|
||
* 3. Await the iframe's resolved Promise (returns t1).
|
||
* 4. Latency = t1 - t0 (ms).
|
||
*/
|
||
async function measureSingleSeek(page: Page, frame: Frame, target: number): Promise<number> {
|
||
await frame.evaluate(
|
||
(target: number, tolerance: number, timeoutMs: number) => {
|
||
window.__perfScrubAwait = new Promise<number>((resolve, reject) => {
|
||
const deadlineWall = performance.timeOrigin + performance.now() + timeoutMs;
|
||
const tick = () => {
|
||
const wall = performance.timeOrigin + performance.now();
|
||
const time = window.__player?.getTime?.() ?? Number.NaN;
|
||
if (Number.isFinite(time) && Math.abs(time - target) < tolerance) {
|
||
resolve(wall);
|
||
return;
|
||
}
|
||
if (wall > deadlineWall) {
|
||
reject(new Error(`[scrub] timeout target=${target} last=${time}`));
|
||
return;
|
||
}
|
||
requestAnimationFrame(tick);
|
||
};
|
||
requestAnimationFrame(tick);
|
||
});
|
||
},
|
||
target,
|
||
MATCH_TOLERANCE_S,
|
||
SEEK_TIMEOUT_MS,
|
||
);
|
||
|
||
const t0Wall = await page.evaluate((targetSeconds: number) => {
|
||
const el = document.getElementById("player") as
|
||
| (HTMLElement & { seek: (t: number) => void })
|
||
| null;
|
||
if (!el) throw new Error("[scenario:scrub] player element missing on host page");
|
||
const wall = performance.timeOrigin + performance.now();
|
||
el.seek(targetSeconds);
|
||
return wall;
|
||
}, target);
|
||
|
||
// Puppeteer awaits the Promise we stashed on window and returns its resolved value.
|
||
const t1Wall = (await frame.evaluate(() => window.__perfScrubAwait as Promise<number>)) as number;
|
||
|
||
return t1Wall - t0Wall;
|
||
}
|
||
|
||
async function runScrubBatch(
|
||
page: Page,
|
||
frame: Frame,
|
||
mode: Mode,
|
||
idx: number,
|
||
total: number,
|
||
): Promise<number[]> {
|
||
const latencies: number[] = [];
|
||
for (const target of SEEK_TARGETS) {
|
||
const latency = await measureSingleSeek(page, frame, target);
|
||
latencies.push(latency);
|
||
}
|
||
const p95 = percentile(latencies, 95);
|
||
console.log(
|
||
`[scenario:scrub] run[${idx + 1}/${total}] mode=${mode} p95=${p95.toFixed(2)}ms n=${latencies.length}`,
|
||
);
|
||
return latencies;
|
||
}
|
||
|
||
async function runOnce(
|
||
opts: ScrubScenarioOpts,
|
||
fixture: string,
|
||
idx: number,
|
||
total: number,
|
||
): Promise<RunResult> {
|
||
const ctx = await opts.browser.createBrowserContext();
|
||
try {
|
||
const page = await ctx.newPage();
|
||
const { duration } = await loadHostPage(page, opts.origin, { fixture });
|
||
const requiredDuration = Math.max(...SEEK_TARGETS);
|
||
if (duration < requiredDuration) {
|
||
throw new Error(
|
||
`[scenario:scrub] fixture composition is ${duration.toFixed(2)}s but scrub targets require >= ${requiredDuration}s`,
|
||
);
|
||
}
|
||
const frame = await getFixtureFrame(page, fixture);
|
||
|
||
// Defensively pause: the host shell doesn't autoplay, but `pause()` also
|
||
// cancels any pending autoplay-on-ready behavior and guarantees the
|
||
// timeline isn't ticking under our seek measurements.
|
||
await page.evaluate(() => {
|
||
const el = document.getElementById("player") as (HTMLElement & { pause?: () => void }) | null;
|
||
el?.pause?.();
|
||
});
|
||
await frame.waitForFunction(() => window.__player?.isPlaying?.() === false, {
|
||
timeout: PAUSE_CONFIRM_TIMEOUT_MS,
|
||
});
|
||
|
||
// Inline mode first — the player's default `_trySyncSeek` path lands the
|
||
// seek synchronously when the iframe is same-origin (which it is here).
|
||
const inlineLatencies = await runScrubBatch(page, frame, "inline", idx, total);
|
||
|
||
// Force isolated mode by shadowing `_trySyncSeek` on the instance with
|
||
// a function that always reports failure. The fallback in `seek()` then
|
||
// sends the seek through `_sendControl("seek", { frame })`, which is the
|
||
// same path a cross-origin embed (or a Phase 1 build without sync seek)
|
||
// would take.
|
||
await page.evaluate(() => {
|
||
const el = document.getElementById("player") as
|
||
| (HTMLElement & { _trySyncSeek?: (t: number) => boolean })
|
||
| null;
|
||
if (!el) throw new Error("[scenario:scrub] player element missing on host page");
|
||
el._trySyncSeek = () => false;
|
||
});
|
||
|
||
const isolatedLatencies = await runScrubBatch(page, frame, "isolated", idx, total);
|
||
|
||
await page.close();
|
||
return { inlineLatencies, isolatedLatencies };
|
||
} finally {
|
||
await ctx.close();
|
||
}
|
||
}
|
||
|
||
export async function runScrub(opts: ScrubScenarioOpts): Promise<Metric[]> {
|
||
const fixture = opts.fixture ?? DEFAULT_FIXTURE;
|
||
const runs = Math.max(1, opts.runs);
|
||
console.log(
|
||
`[scenario:scrub] fixture=${fixture} runs=${runs} seeks_per_mode=${SEEK_TARGETS.length} tolerance=${(MATCH_TOLERANCE_S * 1000).toFixed(0)}ms`,
|
||
);
|
||
|
||
const allInline: number[] = [];
|
||
const allIsolated: number[] = [];
|
||
for (let i = 0; i < runs; i++) {
|
||
const result = await runOnce(opts, fixture, i, runs);
|
||
allInline.push(...result.inlineLatencies);
|
||
allIsolated.push(...result.isolatedLatencies);
|
||
}
|
||
|
||
const inlineP95 = percentile(allInline, 95);
|
||
const isolatedP95 = percentile(allIsolated, 95);
|
||
console.log(
|
||
`[scenario:scrub] aggregate inline_p95=${inlineP95.toFixed(2)}ms isolated_p95=${isolatedP95.toFixed(2)}ms (runs=${runs} samples_per_mode=${allInline.length})`,
|
||
);
|
||
|
||
return [
|
||
{
|
||
name: "scrub_latency_p95_inline_ms",
|
||
baselineKey: "scrubLatencyP95InlineMs",
|
||
value: inlineP95,
|
||
unit: "ms",
|
||
direction: "lower-is-better",
|
||
samples: allInline,
|
||
},
|
||
{
|
||
name: "scrub_latency_p95_isolated_ms",
|
||
baselineKey: "scrubLatencyP95IsolatedMs",
|
||
value: isolatedP95,
|
||
unit: "ms",
|
||
direction: "lower-is-better",
|
||
samples: allIsolated,
|
||
},
|
||
];
|
||
}
|