Files
hyperframes/skills/media-use/scripts/lib/ltx-video-provider.test.mjs
T
James RussoandClaude Opus 5 8392e84a18 fix(media-use): repoint the dead videogen tier, demote past unusable local models (#3509)
* fix(media-use): repoint the dead videogen tier, demote past unusable models

`LOCAL_MODELS.videogen`'s `large` tier named `dgrauet/ltx-2.3-mlx-bf16`, which
returns HTTP 401 and cannot be downloaded at all. It was not a dormant entry:
`rankedByPreference` sorts by descending `needs.ramMB` when no `rank` is set,
so the largest fitting tier is tried FIRST by design. Any machine clearing
32 GB *available* RAM selected the dead entry, `ltxVideoGenerate` caught the
failure and returned a bare `null`, and since `ltx.local` is last in
`["heygen.video", "ltx.local"]` and network providers are skipped under
`--local-only` (`registry.mjs:206`), local video generation failed outright
instead of falling back to the tier that works.

It survived review because the table landed with "live verification on a 24GB
M-series Mac" - and a 24 GB machine cannot select a 32 GB tier, so that entry
was unreachable on the only machine that validated it. The unit fixtures
inherit the same ceiling (`fittingSpecs` is 20000MB), so every existing test
exercised the medium tier alone.

Two changes:

1. Repoint to `dgrauet/ltx-2.3-mlx-q8` (reachable) and correct `sizeMB` from
   45000 to 28800. Measured against the HF API: the q8 repo totals 87.5 GB,
   and the registry's own targeted `--include` subset is 28.76 GB. That
   matches the sibling q4 entry's convention (`sizeMB: 20000` vs a measured
   19.48 GB subset), so 45000 was wrong under either reading. `--low-ram` is
   added because the entry's own note calls it required at this tier's 32 GB
   floor, and the invoke omitted it.

2. A repoint alone is one bad URL from a repeat, so add the missing recovery.
   `selectModelLadder` returns every fitting model best-first;
   `selectModel`'s pick is now defined as that list's head. All three sites
   that previously selected exactly one model and failed terminally walk the
   ladder instead, demoting past a tier that cannot run here - gated weights,
   runner off PATH, an OOM at a tier that nominally fits:

   - `ltx-video-provider.mjs` (videogen, the reported failure)
   - `mflux-provider.mjs` (imagegen - same shape, and its 32 GB/64 GB tiers
     are equally unverifiable on a 24 GB machine)
   - `local-run.mjs` (tts/asr/upscale - `fish-speech` missing should still
     get you Kokoro)

   Every demotion is logged rather than silent, so a quietly smaller model is
   never mistaken for the tier the machine nominally qualified for.

Also fixes the `install` string both videogen entries share: it ended at
`uv sync --all-extras`, which leaves the entry point in `.venv/bin`, so the
"`ltx-2-mlx` not on PATH" hint named a command that following the instruction
would not put on PATH.

The q8 tier is NOT live-verified - no 32 GB+ Apple Silicon machine was
available - and its notes say so. Shipping it unverified is safe precisely
because of change 2: a wrong tier now costs one failed attempt, not the whole
local path.

- Rames Jusso

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(media-use): report the real videogen download size, disclose it, discard failed partials

Addresses review feedback on #3509 (CHANGES_REQUESTED at 90df164a), plus the
follow-on ask to tell the user what a download costs before they accept it.

1. `sizeMB` described a targeted `--include` subset that no run ever gets.
   Both videogen invokes pass a repo id to `--model`, and upstream
   `resolve_model_dir()` (`ltx_pipelines_mlx/utils/_orchestration.py:35-40`)
   calls `snapshot_download(repo)` with no `allow_patterns`, so the full repo
   lands regardless of what was pre-fetched. Corrected to measured repo
   totals: q8 87500 (87,511,991,375 B) and q4 59700 (59,686,429,583 B). q4 was
   wrong the same way at 20000, so both are fixed together rather than leaving
   one convention on each side.

   My earlier claim that 28800 "matches the sibling q4 entry's convention" was
   wrong in the way that matters: the convention itself described a subset the
   runner does not honor. The file's own comment already said "blind
   snapshot-downloads the lot (60 GB q4, 88 GB q8)" three lines above the
   fields that contradicted it, and the original report measured it too ("the
   q4 cache ended at 56 GB and q8 at 82 GB"), which reconciles exactly once
   read as GiB: 59.69 GB = 55.6 GiB, 87.51 GB = 81.5 GiB. So the download is
   the complete repo both times, not a partial fetch.

   Removed the `--include` recipe rather than repairing it: it is ineffective
   (the runner refetches at generate time) and insufficient (`--two-stage` is
   "dev model + CFG at half-res, upscale, distilled LoRA refine" per upstream's
   own help text, so it needs transformer-dev AND transformer-distilled AND
   spatial_upscaler_x2; `--distilled` needs an upscaler too). The q4 tier
   verified on a 24 GB Mac only worked BECAUSE the download is unfiltered.

2. Nothing told the user what they were agreeing to before a tool started
   pulling tens of GB. `describeDownload()` in `specs.mjs` names the size and
   the directory the weights land in, and checks free space with `statfs`
   against that directory rather than cwd, since the weights do not land in
   cwd. A tier that will not fit is still offered, with a plain statement that
   it will not fit: hiding it would make a machine that could free up space
   look like it has no large tier. Unknown free space reports as unknown, not
   as zero. Wired into both providers' install hints, the `runLocalModel`
   install payload (now carrying `sizeMB`), and `describeModelLadder`.

3. Each retry attempt mints its own timestamped temp path, so a partial
   artifact from a failed tier was orphaned rather than overwritten, and a
   lower tier then succeeding hid it. Both providers discard the partial before
   demoting, guarded so a file that cannot be removed never masks the generate
   failure it came from. Video is the material case: a partial mp4 is large.

   `local-run.mjs` is deliberately unchanged here. Its `out` is caller-provided
   and identical across attempts, so a partial is overwritten rather than
   orphaned, and unlinking a path the caller named would be a footgun. The rule
   the two providers follow is: clean up what you allocate.

Tests: 553/553 across `skills/**/*.test.mjs` (+15). New coverage pins the
cleanup (failed tier's partial removed, returned artifact survives, one discard
per attempt on the all-fail path, an unremovable partial still surfaces the
real failure) and the disclosure (cache-dir precedence, statfs walk-up to the
deepest existing ancestor, unknown-vs-zero, and the will-not-fit wording).
Every new guard mutation-tested: removing any one of them turns tests red.

- Rames Jusso

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 16:15:51 -07:00

317 lines
9.7 KiB
JavaScript

import { test } from "node:test";
import assert from "node:assert/strict";
import { dirname } from "node:path";
import { tmpdir } from "node:os";
import { ltxVideoGenerate } from "./ltx-video-provider.mjs";
const fittingSpecs = { availableRamMB: 20000, gpu: { present: true } };
test("no fitting local model: falls through without checking for a binary", async (t) => {
t.mock.method(console, "error", () => {});
const calls = [];
const result = await ltxVideoGenerate(
"a calm ocean wave at sunset",
{ specs: { availableRamMB: 100, gpu: { present: true } } },
(...call) => calls.push(call),
() => true,
);
assert.equal(result, null);
assert.deepEqual(calls, []);
});
test("binary missing from PATH: prints the model install hint and falls through", async (t) => {
const errors = [];
t.mock.method(console, "error", (message) => errors.push(message));
const calls = [];
const fakeExec = (...call) => {
calls.push(call);
throw new Error("not found");
};
const result = await ltxVideoGenerate(
"a calm ocean wave at sunset",
{ specs: fittingSpecs },
fakeExec,
);
assert.equal(result, null);
assert.equal(calls.length, 1);
assert.deepEqual(calls[0].slice(0, 2), ["which", ["ltx-2-mlx"]]);
assert.equal(errors.length, 1);
assert.match(errors[0], /git clone https:\/\/github\.com\/dgrauet\/ltx-2-mlx/);
// the install hint is the accept moment: say what the pull costs
assert.match(errors[0], /GB of weights to/);
});
test("generate argv substitutes a spaced prompt after tokenizing and uses verified defaults", async () => {
const calls = [];
const checkedPaths = [];
const fakeExec = (...call) => calls.push(call);
const pathExists = (path) => {
checkedPaths.push(path);
return false;
};
const intent = "a calm ocean wave at sunset";
const result = await ltxVideoGenerate(intent, { specs: fittingSpecs }, fakeExec, pathExists);
assert.equal(result, null);
assert.equal(calls.length, 2);
const [bin, argv, opts] = calls[1];
assert.equal(bin, "ltx-2-mlx");
assert.equal(opts.timeout, 1_800_000);
const expectedPairs = [
["--prompt", intent],
["--width", "512"],
["--height", "320"],
["--frames", "33"],
["--output", checkedPaths[0]],
];
let previousIndex = -1;
for (const [flag, value] of expectedPairs) {
const index = argv.indexOf(flag);
assert.ok(index > previousIndex, `${flag} should follow the previous required option`);
assert.equal(argv[index + 1], value);
previousIndex = index;
}
assert.equal(argv.filter((arg) => arg === intent).length, 1);
});
test("successful generation returns the generated MP4 result", async () => {
const calls = [];
const fakeExec = (...call) => calls.push(call);
const intent = "a calm ocean wave at sunset";
const result = await ltxVideoGenerate(intent, { specs: fittingSpecs }, fakeExec, () => true);
assert.ok(result);
assert.equal(calls.length, 2);
assert.equal(dirname(result.localPath), tmpdir());
assert.match(result.localPath, /media-use-ltx-\d+-\d+\.mp4$/);
assert.deepEqual(result, {
localPath: result.localPath,
ext: ".mp4",
source: "generated",
metadata: {
description: intent,
provider: "ltx.local",
provenance: { prompt: intent },
},
});
});
test("generate failure returns null instead of throwing", async (t) => {
t.mock.method(console, "error", () => {});
let calls = 0;
const fakeExec = () => {
calls += 1;
if (calls === 2) {
const error = new Error("generation failed");
error.stderr = "LTX failed";
throw error;
}
};
const result = await ltxVideoGenerate(
"storm clouds",
{ specs: fittingSpecs },
fakeExec,
() => true,
);
assert.equal(result, null);
assert.equal(calls, 2);
});
test("missing generated output returns null and says so", async (t) => {
const errors = [];
t.mock.method(console, "error", (message) => errors.push(message));
const result = await ltxVideoGenerate(
"storm clouds",
{ specs: fittingSpecs },
() => {},
() => false,
);
assert.equal(result, null);
assert.equal(errors.length, 1);
assert.match(errors[0], /wrote no output file/);
});
// 40GB clears BOTH videogen tiers, so the ladder has two rungs. `fittingSpecs`
// above sits under the large tier's floor on purpose: every other test in this
// file exercises the medium tier alone, which is precisely why a broken large
// tier could sit in the table unnoticed.
const bothTiersSpecs = { availableRamMB: 40000, gpu: { present: true } };
const isGenerate = (call) => call[0] !== "which";
test("a top tier that cannot run demotes to the next fitting tier", async (t) => {
const errors = [];
t.mock.method(console, "error", (message) => errors.push(message));
const calls = [];
// The runner is installed, but the large tier's weights are gated: the
// download 401s and `generate` exits non-zero. The medium tier then works.
const fakeExec = (...call) => {
calls.push(call);
if (isGenerate(call) && call[1].includes("dgrauet/ltx-2.3-mlx-q8")) {
const err = new Error("exit 1");
err.stderr = "401 Client Error: Unauthorized for url: .../ltx-2.3-mlx-q8";
throw err;
}
};
const result = await ltxVideoGenerate(
"storm clouds",
{ specs: bothTiersSpecs },
fakeExec,
() => true,
);
assert.ok(result, "the medium tier still produced a video");
const generated = calls.filter(isGenerate).map((call) => call[1].join(" "));
assert.equal(generated.length, 2, "large attempted first, then medium");
assert.match(generated[0], /dgrauet\/ltx-2\.3-mlx-q8/);
assert.match(generated[1], /dgrauet\/ltx-2\.3-mlx-q4/);
// the demotion is reported, never silent: a smaller model changes the output
assert.equal(errors.length, 1);
assert.match(errors[0], /ltx-2\.3-mlx-q8\) failed/);
assert.match(errors[0], /401/);
});
test("every fitting tier failing returns null, one reason per tier", async (t) => {
const errors = [];
t.mock.method(console, "error", (message) => errors.push(message));
const fakeExec = (...call) => {
if (isGenerate(call)) throw new Error("mlx out of memory");
};
const result = await ltxVideoGenerate(
"storm clouds",
{ specs: bothTiersSpecs },
fakeExec,
() => true,
);
assert.equal(result, null);
assert.equal(errors.length, 2, "both tiers tried, both reported");
assert.match(errors[0], /ltx-2\.3-mlx-q8/);
assert.match(errors[1], /ltx-2\.3-mlx-q4/);
});
test("preferTier pins the attempt to one tier instead of demoting", async (t) => {
t.mock.method(console, "error", () => {});
const calls = [];
const fakeExec = (...call) => {
calls.push(call);
if (isGenerate(call)) throw new Error("boom");
};
const result = await ltxVideoGenerate(
"storm clouds",
{ specs: bothTiersSpecs, preferTier: "large" },
fakeExec,
() => true,
);
assert.equal(result, null);
const generated = calls.filter(isGenerate).map((call) => call[1].join(" "));
assert.equal(generated.length, 1, "pinned to large: no demotion to medium");
assert.match(generated[0], /dgrauet\/ltx-2\.3-mlx-q8/);
});
// A failed attempt's temp path is minted per attempt (it carries a timestamp),
// so without cleanup a partial mp4 from a failed tier is orphaned rather than
// overwritten - and a lower tier then succeeding hides it. Partial video files
// are the expensive case, which is why this is pinned.
const outputOf = (argv) => argv[argv.indexOf("--output") + 1];
test("a failed attempt's partial output is discarded before demoting", async (t) => {
t.mock.method(console, "error", () => {});
const unlinked = [];
const attempted = [];
const fakeExec = (...call) => {
if (!isGenerate(call)) return;
attempted.push(outputOf(call[1]));
if (call[1].includes("dgrauet/ltx-2.3-mlx-q8")) {
// OOM mid-write is one of the advertised demotion cases
const err = new Error("exit 1");
err.stderr = "mlx.core.metal: out of memory";
throw err;
}
};
const result = await ltxVideoGenerate(
"storm clouds",
{ specs: bothTiersSpecs },
fakeExec,
() => true,
(path) => unlinked.push(path),
);
assert.ok(result, "the medium tier still produced a video");
assert.deepEqual(unlinked, [attempted[0]], "the failed large-tier partial is removed");
});
test("every tier failing discards every partial, one per attempt", async (t) => {
t.mock.method(console, "error", () => {});
const unlinked = [];
const attempted = [];
const fakeExec = (...call) => {
if (!isGenerate(call)) return;
attempted.push(outputOf(call[1]));
throw new Error("mlx out of memory");
};
const result = await ltxVideoGenerate(
"storm clouds",
{ specs: bothTiersSpecs },
fakeExec,
() => true,
(path) => unlinked.push(path),
);
assert.equal(result, null);
assert.equal(attempted.length, 2, "both tiers attempted");
assert.deepEqual(unlinked, attempted, "nothing is left behind on the all-fail path");
});
test("a successful generation is never discarded", async () => {
const unlinked = [];
const result = await ltxVideoGenerate(
"storm clouds",
{ specs: bothTiersSpecs },
() => {},
() => true,
(path) => unlinked.push(path),
);
assert.ok(result);
assert.deepEqual(unlinked, [], "the returned artifact must survive");
});
test("an unremovable partial does not mask the generate failure", async (t) => {
t.mock.method(console, "error", () => {});
const fakeExec = (...call) => {
if (isGenerate(call)) throw new Error("mlx out of memory");
};
const result = await ltxVideoGenerate(
"storm clouds",
{ specs: bothTiersSpecs },
fakeExec,
() => true,
() => {
throw new Error("EPERM: operation not permitted");
},
);
// cleanup is best-effort: a partial we cannot delete must not become the error
assert.equal(result, null);
});