feat(media-use): add video generation (HeyGen avatar-video + local LTX fallback) (#2614)

* fix(media-use): tag HeyGen TTS generation with attribution header

Centralizes the X-HeyGen-Client-Source header into HEYGEN_CLIENT_SOURCE_ARGV
in heygen-cli.mjs and reuses it in heygen-search.mjs (dropping the duplicated
inline literal) so voice-provider's `voice speech create` call carries it too.
The generation call was previously untagged, making media-use TTS usage
invisible in HeyGen's billing/analytics warehouse; the read-only `voice list`
discovery call intentionally stays untagged.

* feat(media-use): add local LTX video generate provider

* feat(media-use): add HeyGen avatar-video generate provider

* feat(media-use): register video as a real provider type

* docs(media-use): document the wired video type and full HeyGen tagging coverage

resolve --type video is now the default path (HeyGen avatar video first,
local LTX fallback, sign-in nudge on auth failure) instead of a manual
recipe; correct the claim that only search requests are tagged now that
TTS and avatar-video generation carry the attribution header too.

* fix(media-use): wire --avatar-id/--voice-id CLI flags and close video-provider auth/cache gaps

- resolve.mjs never implemented the --avatar-id/--voice-id override that
  operations.md documented, so following the docs crashed with
  ERR_PARSE_ARGS_UNKNOWN_OPTION; wire the flags through to ctx.
- defaultAvatarId/defaultStarfishVoiceId cached a failed discovery lookup
  as a permanent null, disabling heygen.video after one transient miss;
  cache only a truthy id, matching the same fix in voice-provider.mjs's
  defaultVoiceId.
- the avatar-video onboarding nudge only fired on a video-create failure,
  never when avatar/voice discovery itself was unauthenticated (the
  common unauthenticated case) -- propagate the discovery failure reason
  so onboarding fires either way.
- dedupe the CLI-shelling JSON helper (heygen-cli.mjs's new runHeygenJson)
  and the local-model argv-template builder (local-models.mjs's new
  buildArgv) instead of leaving byte-identical copies in each provider.

* fix(media-use): address avatar-video PR review feedback

- heygenVideoGenerate short-circuits after the first discovery-call
  failure instead of always attempting both avatar list and voice list,
  so an unauthenticated caller gets one onboarding message and one
  provider-error telemetry ping instead of a double-fire.
- runHeygenJson logs a diagnostic when a CLI call succeeds but returns
  unparseable JSON, instead of silently returning null.
- dedupe the "avatar video is free" onboarding string into one constant
  (was duplicated across three call sites).

* fix(media-use): match review-requested naming and message conventions

- export AVATAR_VIDEO_SIGNIN_MESSAGE from heygen-video-provider.mjs so
  the test imports the canonical string instead of redeclaring it.
- runHeygenJson's non-JSON diagnostic now matches heygen-search.mjs's
  existing wording ("returned non-JSON output").
This commit is contained in:
Miguel Ángel
2026-07-17 03:56:57 -04:00
committed by GitHub
parent f084a7217d
commit 0a66671fc5
18 changed files with 1021 additions and 122 deletions
+17 -26
View File
@@ -1,29 +1,9 @@
import { execFileSync } from "node:child_process";
import { reportHeygenFailure } from "./heygen-cli.mjs";
import { HEYGEN_CLIENT_SOURCE_ARGV, runHeygenJson } from "./heygen-cli.mjs";
// Voice / TTS generation via the HeyGen CLI — the only external CLI media-use
// shells (CLI-only invariant: media-use holds no keys; the CLI owns auth).
// Flags verified against `heygen voice speech create --help` (v0.3.0).
function runJson(bin, argv, label) {
let out;
try {
out = execFileSync(bin, argv, {
encoding: "utf8",
timeout: 120000,
stdio: ["pipe", "pipe", "pipe"],
});
} catch (err) {
reportHeygenFailure(err, `${bin} ${label}`);
return null;
}
try {
return JSON.parse(out);
} catch {
return null;
}
}
function result(url, duration, provider, intent) {
if (!url) return null;
return {
@@ -41,11 +21,13 @@ function result(url, duration, provider, intent) {
// HeyGen TTS requires a starfish-engine voice. Default to the first one the
// catalog returns (deterministic order); pass ctx.voiceId to override.
// ponytail: listed once per process; the resolved asset is frozen + cached after
// first use, so the network list only happens on a cache miss.
// first use, so the network list only happens on a cache miss. Cache only a
// truthy id -- a transient list failure must not poison the cache with `null`
// and permanently disable TTS for the rest of the process.
let cachedVoiceId;
function defaultVoiceId() {
if (cachedVoiceId !== undefined) return cachedVoiceId;
const j = runJson(
if (cachedVoiceId) return cachedVoiceId;
const j = runHeygenJson(
"heygen",
["voice", "list", "--engine", "starfish", "--limit", "1"],
"voice list",
@@ -57,9 +39,18 @@ function defaultVoiceId() {
export async function heygenTtsGenerate(intent, ctx) {
const voiceId = ctx?.voiceId || defaultVoiceId();
if (!voiceId) return null;
const p = runJson(
const p = runHeygenJson(
"heygen",
["voice", "speech", "create", "--text", intent, "--voice-id", voiceId],
[
...HEYGEN_CLIENT_SOURCE_ARGV,
"voice",
"speech",
"create",
"--text",
intent,
"--voice-id",
voiceId,
],
"tts",
);
return result(p?.data?.audio_url, p?.data?.duration, "heygen.tts", intent);