perf(producer): hdr benchmark harness — --tags filter, peak heap/RSS tracking, bench:hdr script (#382)

## Summary

Make the existing benchmark harness genuinely useful for HDR perf work: positive `--tags` filter, peak heap/RSS sampling, a `bench:hdr` script, and a perf README documenting the captured April-2026 baseline. Lands first in the Chunk 8 sub-stack so subsequent perf PRs can be measured against a known starting point.

## Why

`Chunk 8A` of `plans/hdr-followups.md`. Wall-clock timing alone can't catch slow memory regressions like an unbounded image cache — peak RSS does. And the existing harness only had `--exclude-tags`, so HDR runs had to wait for unrelated SDR fixtures.

## What changed

**1. Positive `--tags` filter** in `benchmark.ts`. Adds `--tags hdr` so HDR runs don't have to wait for unrelated fixtures. Filters compose: a fixture must match `--tags` (if provided) AND must not match `--exclude-tags`.

**2. Peak heap + RSS tracking** in `executeRenderJob`. A 250 ms periodic `process.memoryUsage()` sampler runs alongside every render and reports `peakRssMb` / `peakHeapUsedMb` in `RenderPerfSummary`. Sampler is `unref`'d and always cleared in `finally` so it never keeps the event loop alive or leaks across jobs. Both fields are optional on the interface for back-compat with serialized older summaries.

**3. `bench:hdr` convenience script** plus a perf README at `tests/perf/README.md` documenting the harness, the new flags, and the captured April-2026 HDR baseline (PQ regression: 34.5 s / 272 MiB RSS, HLG regression: 11.5 s / 227 MiB RSS, both 1080p / 1 worker / 1 run).

The benchmark output table is widened and gains `PeakRSS` / `PeakHeap` columns. A new `avgOrNull` helper preserves `null` in the JSON when no run reported memory (avoids silently coercing missing data to 0 in older snapshots).

No behavior change for non-benchmark renders — the sampler runs in every `executeRenderJob` but its overhead is a single `process.memoryUsage()` call every 250 ms, well below noise.

## Test plan

- [x] `bunx tsc --noEmit -p packages/producer` — clean.
- [x] `bunx oxlint` / `bunx oxfmt --check` on changed files — clean.
- [x] `bun test src/services/` — 60/60 pass (frameDirCache, orchestrator, etc.).
- [x] `bunx tsx src/benchmark.ts --tags hdr --runs 1` — both HDR fixtures render successfully, summary table prints `PeakRSS`/`PeakHeap` columns, per-run output shows new memory line.
- [x] `bunx tsx src/benchmark.ts --tags nonexistent` — exits 1 with a helpful message naming the active filters.

## Stack

Chunk 8A of `plans/hdr-followups.md`. First PR in the Chunk 8 perf sub-stack; subsequent PRs (image cache, logger gating) measured against this baseline.
This commit is contained in:
Vance Ingalls
2026-04-23 14:49:49 -07:00
committed by GitHub
parent cc9403b6bd
commit 3da8c2e969
4 changed files with 224 additions and 23 deletions
+100
View File
@@ -0,0 +1,100 @@
# Producer Performance Benchmarks
End-to-end render benchmark harness driven by `src/benchmark.ts`. Discovers
fixtures under `packages/producer/tests/<fixture>/` (any directory with a
`meta.json`), runs them through the full producer pipeline, and emits per-stage
timing plus peak memory metrics into `benchmark-results.json`.
The harness is deliberately lightweight — it doesn't enforce thresholds. It's
designed for **regression spotting**: capture a baseline, change something,
re-run, eyeball the diff. For pass/fail thresholds see `tests/perf/baseline.json`
and the perf-regression checks in the integration test suite.
## Quick start
```bash
# Run every non-slow fixture once
cd packages/producer
bun run benchmark
# HDR-only baseline (PQ + HLG fixtures, ~50s on M-series Macs)
bun run bench:hdr
# Average a fixture across multiple runs
bunx tsx src/benchmark.ts --tags hdr --runs 3
# Just the PQ regression
bunx tsx src/benchmark.ts --only hdr-regression
# Skip slow fixtures explicitly (default behavior; here for clarity)
bunx tsx src/benchmark.ts --exclude-tags slow
```
Results are written to
`packages/producer/tests/perf/benchmark-results.json` and a summary table is
printed to stdout.
## CLI flags
| Flag | Description |
| --- | --- |
| `--runs N` | Run each fixture `N` times and average (default: 1). |
| `--only <id>` | Run a single fixture by directory name. |
| `--tags a,b` | **Positive** filter: only fixtures whose `meta.json#tags` contains *any* of the listed tags. |
| `--exclude-tags a,b` | **Negative** filter: skip fixtures with any matching tag. Defaults to `slow`. |
`--tags` and `--exclude-tags` apply independently — a fixture must match the
positive filter (if any) **and** must not match the negative filter.
## Reading the output
Each fixture row prints averaged stage timings plus peak memory:
```
Fixture Total Compile Extract Audio Capture Encode PeakRSS PeakHeap
hdr-hlg-regression 11549ms 187ms 520ms 36ms 8373ms 2394ms 227MiB 69MiB
hdr-regression 34452ms 94ms 1268ms 48ms 27034ms 5914ms 272MiB 118MiB
```
- **Total** — wall-clock time from job submission to mux-complete.
- **Capture** — frame extraction + composition + alpha blit (HDR path).
- **Encode** — chunked or streaming HDR encoder time (HEVC Main10 for HDR).
- **PeakRSS / PeakHeap** — sampled every 250ms inside `executeRenderJob` from
`process.memoryUsage()`; surfaces gross memory regressions (e.g. unbounded
image-cache growth) that wall-clock numbers miss. RSS includes native
ffmpeg/Chrome allocations; heap is JS-side V8 only.
## HDR baseline (April 2026)
Captured on macOS arm64 (M-series), Bun runtime, 1 worker, default config,
single run. These are illustrative — re-baseline locally before comparing your
own runs.
| Fixture | Total | Capture | Encode | PeakRSS | PeakHeap | Notes |
| --- | ---: | ---: | ---: | ---: | ---: | --- |
| `hdr-hlg-regression` | 11.5s | 8.4s (72%) | 2.4s (21%) | 227 MiB | 69 MiB | 150 frames, 2 HLG sources |
| `hdr-regression` | 34.5s | 27.0s (78%) | 5.9s (17%) | 272 MiB | 118 MiB | 600 frames, 9 PQ sources, shader transition |
Capture dominates HDR runs (~72-78%). The second-biggest cost is HEVC Main10
encode. Memory peaks scale with source count and resolution — the PQ
regression's nine HDR sources push heap from ~70 MiB → ~120 MiB.
When evaluating an HDR optimization (image cache, gated debug logging, etc.)
the metric to watch first is **Capture** ms-per-frame:
```
hdr-regression: capture avg 45ms/frame
hdr-hlg-regression: capture avg 56ms/frame
```
## When to re-baseline
- After landing any change that touches `renderOrchestrator.ts`,
`streamingEncoder.ts`, the HDR alpha-blit path, or `frameDirCache.ts`.
- Before opening a PR that claims a perf win — paste before/after numbers in
the PR description.
- Quarterly, even without code changes, to track infra/dependency drift.
The `bench:hdr` script is the recommended command for routine HDR perf checks
because it filters out non-HDR fixtures (which can be 10× slower without
contributing signal to HDR-specific work).