Files
hyperframes/packages/cli/src/commands/grade-compare.test.ts
T
Miguel ÁngelandClaude Opus 4.8 57b3c78987 feat(media-use): color grading — grade/lut resolve, smart-grade, grade-compare + compare (#2041)
* feat(media-use): color grading — grade/lut resolve, smart-grade, grade-compare CLI

Add color grading to media-use as first-class resolve types plus a faithful
comparison command. All local, offline, deterministic — no model, no GPU.

- resolve -t grade / -t lut: produce a data-color-grading block (or a frozen
  .cube). Look cascade: core preset (no file) -> bundled .cube library ->
  parametric buildCube. Emitted .cube is Rec.709 and validated against core's
  colorLuts constraints (LUT_3D_SIZE <= 64) before it is frozen.
- smart grade (grade --for <media>): ffmpeg signalstats -> adjust suggestion
  (exposure / contrast / white balance), surfaced with the measured evidence on
  stderr as a starting point; never auto-applied.
- hyperframes grade-compare: renders N candidate grades onto a reference frame
  through the real runtime shader into one labeled comparison PNG, so an agent
  picks a look without opening Studio. Prepends an "original" baseline cell by
  default (--no-baseline to omit). Shares the headless-capture pipeline with
  snapshot via capture/captureCompositionFrame.
- media-use SKILL: proactive "media opportunity pass" guidance (grounded
  signal -> offer, ask once, surface don't mutate).

Verified: media-use 116/116, grade-compare 7/7, snapshot 9/9, lint + format
clean, full build green, comparison renders end to end.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* test(cli): narrow grade-compare baseline assertion off unknown-typed grading

Assert the whole cell via toEqual instead of reaching into .grading.preset /
.grading.lut on the unknown-typed field, keeping the test typecheck-clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* feat(media-use): agent-authored LUTs via --params + validate --from cube; never-read-.cube guardrail

- resolve -t lut / -t grade --params '<json>': build a parametric .cube from
  explicit params (bypassing the intent cascade), validate, and freeze in one
  step. --intent becomes the optional description. Lets an agent commit a look
  it computed itself.
- --from <file.cube> now validates the ingested LUT for lut/grade types and
  rejects an invalid/oversized cube (no partial write) — the escape hatch for a
  LUT the agent generated with its own code.
- SKILL.md: hard rule to never read a .cube body into context (~size^3 lines,
  zero legible signal) — inspect via grade-compare (see it) or cube-validate
  (ok/size), read the manifest description for meaning; plus both authoring
  paths and the parametric-vs-film-stock ceiling note.

Verified: media-use 116/116, lint + format clean; smokes — --params builds a
valid frozen cube, grade --params returns a lut block, bad JSON and an oversized
--from cube are both rejected with no stray file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* fix(cli): grade-compare validates referenced LUTs, warns on no-op cells, caps candidates

Bug-bash follow-ups — grade-compare silently accepted bad input:

- Validate LUT *content*, not just existence: each referenced .cube is parsed
  with core's parseCubeLut (now exported from @hyperframes/core) and rejected
  with a per-cell error ("LUT for \"<label>\" is not a valid .cube: ..."). A
  file that exists but isn't a valid cube no longer renders a silent no-op cell.
- Warn on inactive cells: a grading that normalizes to inactive (e.g. a
  malformed {lut:12345}) emits a stderr warning naming the cell; the
  auto-prepended "original" baseline is intentionally inactive and stays silent.
  stdout remains valid JSON.
- Cap candidates at 16 (excluding baseline): over-cap input renders the first N
  and reports {truncated:true, total:M} on stdout + a stderr note — no silent
  drop, no unbounded giant sheet.

Verified: grade-compare 10/10; non-cube LUT → clear error; {lut:12345} → warning
+ ok; 20 cells → cells=17 truncated total=20; valid runs unchanged. Lint/format
clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* feat(cli): general `hyperframes compare` visual-variant primitive

Generalize grade-compare's "render N variants → one labeled sheet → the agent
looks and picks" loop into a standalone command that works on ANY variation
(font, layout, motion, grade, whole compositions) — the tool never needs to
know what differs.

- `hyperframes compare <path...> [--at <sec>] [--labels a,b,c] [--out] [--cols]
  [--json]`: renders each agent-authored composition variant through the real
  runtime (captureCompositionFrame) and stitches one labeled comparison sheet +
  JSON ({ok, sheet, rendered, variants, truncated?/total?}). 2+ paths required;
  caps at 16 with loud truncation. It presents, it does not judge — choosing is
  the caller's job.
- Factored the shared "render a labeled set → contact sheet" path so compare,
  grade-compare, and snapshot all sit on it (no duplication). grade-compare is
  now the first color-specific specialization of this primitive.
- New pathArgs util + contactSheet test; hyperframes-cli SKILL documents compare
  as the agent's "see your own renders and choose" primitive.

Verified: 26/26 across compare + grade-compare + snapshot + contactSheet (no
regressions); compare renders 3 variants into one visibly-distinct labeled
sheet; 2+-path error path clean; lint/format clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* fix(ci): green the skills CI — skip ffmpeg tests when absent, oxfmt markdown

The "Test: skills" CI job runs bare `node --test` with no ffmpeg on PATH (by
design — skills tests are meant to be node-builtin-only). The grade-analyzer +
smart-grade tests shell to ffmpeg and were failing there with ENOENT. Guard
them to skip when ffmpeg isn't on PATH; they still run locally / where it is.

Also oxfmt README.md + hyperframes/media-use SKILL.md (the whole-repo
`oxfmt --check .` Format job caught markdown left unformatted by the rebase
conflict resolution).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* fix(ci): skip core-conformance test when tsx is unavailable

The "Test: skills" CI job installs no deps, so the normalizeHfColorGrading
conformance test (which imports core's TS via `node --import tsx`) failed there.
Guard it to skip when tsx can't resolve; runs locally / in the deps-installed
Test job. Completes the skills-CI greening (the ffmpeg guards handled the rest).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* fix(cli): escape grade-compare src double-quotes (CodeQL XSS) + Windows-safe compare test

- grade-compare built `<img src="...">` (double-quoted) with the single-quote
  escaper, leaving `"` unescaped — a `"` in the frame path could break out
  (CodeQL: incomplete HTML attribute sanitization). Use escapeXml for src.
- compare label test hard-coded POSIX paths that can't match on Windows; assert
  the derived labels (the subject); path resolution is covered elsewhere.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* refactor(media-use): generate LUT library from params (drop committed .cube files)

The 3 bundled .cube files were 733 lines each (2,199 total) and were themselves
buildCube output — pure repo bloat. Replace with compact per-look params in
luts/index.json, generated on resolve; add an optional `url` for future scanned
LUTs to be CDN-hosted + downloaded on demand (freezeUrl) instead of committed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* feat(media-use): serve library LUTs from CDN on-demand (static.heygen.ai/luts), params fallback

Looks now carry a CDN `url` (hosted at s3://heygen-public/luts → static.heygen.ai/luts/<id>.cube);
resolve downloads + validates + freezes on demand, like bgm/image. `params` stays
as the deterministic offline fallback (--local-only, or if the download fails), so
resolution is never blocked on the network. Provider prefers url, falls back to params.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

* fix(media-use): address #2041 review — atomic LUT writes, compare telemetry, follow-ups

- Atomic .cube writes: library provider (url + params) and the parametric
  generator now write to a .tmp path, validate, then rename, so a crash can
  never orphan an invalid .cube at the final path (was validate-after-write).
- track("media_use_resolve") now emits provenance.via (url/params-fallback/params).
- grade-compare + compare: --timeout flag (was hardcoded 5000) and a
  media_use_compare event (cells, truncated, total, render_ready_timed_out);
  openSettledCompositionPage now surfaces the render-ready timeout.
- compare staging skips node_modules/.git; --for gets an upfront existence check.
- Rec.709 luma comment; HYPERFRAMES_ANALYZE_TIMEOUT_MS override; measured note
  uses basename; LUT s3 hosting moved from index.json into luts/README.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H5k87mPZ4d6yiFwcWSb8Vv

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-08 22:20:16 -04:00

254 lines
8.6 KiB
TypeScript

import { readFileSync, writeFileSync } from "node:fs";
import { mkdtempSync, rmSync } from "node:fs";
import { tmpdir } from "node:os";
import { basename, join } from "node:path";
import { HF_COLOR_GRADING_ATTR, serializeHfColorGrading } from "@hyperframes/core";
import { describe, expect, it, vi } from "vitest";
import {
buildGradeCompareHtml,
buildGradeCompareSuccessPayload,
capCandidateCells,
parseGradeCompareArgs,
parseGradesFile,
prepareGradeCompareTempProject,
prependBaselineCell,
resolveLutCells,
warnInactiveGradingCells,
} from "./grade-compare.js";
function tempDir(): string {
return mkdtempSync(join(tmpdir(), "hf-grade-compare-test-"));
}
function validCubeLut(): string {
return `LUT_3D_SIZE 2
0 0 0
1 0 0
0 1 0
1 1 0
0 0 1
1 0 1
0 1 1
1 1 1
`;
}
describe("parseGradesFile", () => {
it("parses a valid grades array into labeled cells", () => {
const dir = tempDir();
try {
const file = join(dir, "grades.json");
writeFileSync(
file,
JSON.stringify([
{ label: "warm", grading: { preset: "warm-daylight" } },
{ label: "punch", grading: { adjust: { exposure: 0.5, contrast: 0.4 } } },
]),
);
expect(parseGradesFile(file)).toEqual([
{ label: "warm", grading: { preset: "warm-daylight" } },
{ label: "punch", grading: { adjust: { exposure: 0.5, contrast: 0.4 } } },
]);
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
it("errors clearly for missing or invalid grades files", () => {
const dir = tempDir();
try {
expect(() => parseGradesFile(join(dir, "missing.json"))).toThrow(
/Grades file not found: .*missing\.json/,
);
const invalid = join(dir, "invalid.json");
writeFileSync(invalid, JSON.stringify({ label: "not an array" }));
expect(() => parseGradesFile(invalid)).toThrow(
"Grades file must be a JSON array of { label, grading } objects",
);
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
});
describe("prependBaselineCell", () => {
it("prepends an ungraded 'original' reference cell", () => {
const candidates = resolveLutCells("/tmp/luts/film.cube");
const withBaseline = prependBaselineCell(candidates);
expect(withBaseline).toHaveLength(candidates.length + 1);
// empty grading (no preset / no lut) → normalizes to inactive → renders the
// source frame untouched, giving a reference to judge candidates against
expect(withBaseline.at(0)).toEqual({ label: "original", grading: {} });
expect(withBaseline.slice(1)).toEqual(candidates);
});
});
describe("resolveLutCells", () => {
it("expands comma-separated LUT paths into labeled grading cells", () => {
expect(resolveLutCells("/tmp/luts/film.cube,./cool.look.cube")).toEqual([
{ label: "film", grading: { lut: { src: "/tmp/luts/film.cube" } } },
{ label: "cool.look", grading: { lut: { src: "./cool.look.cube" } } },
]);
});
});
describe("parseGradeCompareArgs", () => {
it("requires exactly one grade source", () => {
expect(() => parseGradeCompareArgs({ for: "frame.png" })).toThrow(
"Exactly one of --grades or --luts is required",
);
expect(() =>
parseGradeCompareArgs({ for: "frame.png", grades: "grades.json", luts: "a.cube" }),
).toThrow("Exactly one of --grades or --luts is required");
});
});
describe("buildGradeCompareHtml", () => {
it("renders one labeled color-graded image per cell with composition metadata", () => {
const cells = [
{ label: "warm <daylight>", grading: { preset: "warm-daylight" } },
{ label: "cool", grading: { adjust: { temperature: -0.8 } } },
{ label: "punchy", grading: { adjust: { exposure: 0.5, contrast: 0.4 } } },
];
const html = buildGradeCompareHtml({
cells,
frameSrc: "frame.png",
frameWidth: 640,
frameHeight: 360,
});
const escapedAttr = HF_COLOR_GRADING_ATTR.replace(/[-/\\^$*+?.()|[\]{}]/g, "\\$&");
const attrPattern = new RegExp(`<img[^>]+${escapedAttr}=`, "g");
expect(html.match(attrPattern)).toHaveLength(cells.length);
expect(html).toContain('data-composition-id="grade-compare"');
expect(html).toContain('data-width="1168"');
expect(html).toContain('data-height="742"');
expect(html).toContain('data-duration="1"');
expect(html).toContain("warm &lt;daylight&gt;");
for (const cell of cells) {
expect(html).toContain(`${HF_COLOR_GRADING_ATTR}='${serializeHfColorGrading(cell.grading)}'`);
}
});
});
describe("prepareGradeCompareTempProject", () => {
it("copies the frame and LUTs into a temp project and rewrites LUT src values", async () => {
const dir = tempDir();
try {
const framePath = join(dir, "frame.png");
const lutPath = join(dir, "look.cube");
writeFileSync(framePath, "fake-png");
writeFileSync(lutPath, validCubeLut());
const prepared = await prepareGradeCompareTempProject({
projectDir: dir,
framePath,
frameBuffer: Buffer.from("fake-png"),
cells: [
{ label: "film", grading: { lut: { src: basename(lutPath), intensity: 0.7 } } },
{ label: "warm", grading: { preset: "warm-daylight" } },
],
frameWidth: 640,
frameHeight: 360,
});
try {
expect(readFileSync(join(prepared.tempDir, "frame.png"), "utf-8")).toBe("fake-png");
expect(readFileSync(join(prepared.tempDir, "lut-0.cube"), "utf-8")).toBe(validCubeLut());
expect(readFileSync(join(prepared.tempDir, "index.html"), "utf-8")).toContain(
`${HF_COLOR_GRADING_ATTR}='${serializeHfColorGrading({
lut: { src: "lut-0.cube", intensity: 0.7 },
})}'`,
);
} finally {
rmSync(prepared.tempDir, { recursive: true, force: true });
}
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
it("rejects an existing LUT file whose .cube content cannot be parsed", async () => {
const dir = tempDir();
try {
const framePath = join(dir, "frame.png");
const lutPath = join(dir, "broken.cube");
writeFileSync(framePath, "fake-png");
writeFileSync(lutPath, "plain text, not cube data\n");
await expect(
prepareGradeCompareTempProject({
projectDir: dir,
framePath,
frameBuffer: Buffer.from("fake-png"),
cells: [{ label: "broken look", grading: { lut: { src: basename(lutPath) } } }],
frameWidth: 640,
frameHeight: 360,
}),
).rejects.toThrow(/LUT for "broken look" is not a valid \.cube:/);
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
});
describe("warnInactiveGradingCells", () => {
it("warns to stderr for normalized-but-inactive candidate grades", () => {
const errorSpy = vi.spyOn(console, "error").mockImplementation(() => {});
const logSpy = vi.spyOn(console, "log").mockImplementation(() => {});
try {
expect(() =>
warnInactiveGradingCells([{ label: "numeric lut", grading: { lut: 12345 } }]),
).not.toThrow();
expect(errorSpy).toHaveBeenCalledWith(
expect.stringContaining(
'Warning: grading for "numeric lut" is inactive/no-op — it will render ungraded',
),
);
expect(logSpy).not.toHaveBeenCalled();
} finally {
vi.restoreAllMocks();
}
});
});
describe("capCandidateCells", () => {
it("truncates over-cap candidates and exposes truncation metadata for JSON output", () => {
const errorSpy = vi.spyOn(console, "error").mockImplementation(() => {});
const logSpy = vi.spyOn(console, "log").mockImplementation(() => {});
try {
const cells = Array.from({ length: 18 }, (_, index) => ({
label: `candidate ${index + 1}`,
grading: { adjust: { exposure: index / 10 } },
}));
const capped = capCandidateCells(cells);
expect(capped.cells).toHaveLength(16);
expect(capped.cells.at(0)?.label).toBe("candidate 1");
expect(capped.cells.at(15)?.label).toBe("candidate 16");
expect(capped.truncated).toBe(true);
expect(capped.total).toBe(18);
expect(buildGradeCompareSuccessPayload("grade-compare.png", 17, capped)).toEqual({
ok: true,
sheet: "grade-compare.png",
cells: 17,
truncated: true,
total: 18,
});
expect(errorSpy).toHaveBeenCalledWith(
expect.stringContaining(
"Warning: 18 candidate grades exceed the 16-cell cap — rendering the first 16 of 18; re-run with fewer grades or split into multiple runs.",
),
);
expect(logSpy).not.toHaveBeenCalled();
} finally {
vi.restoreAllMocks();
}
});
});