mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-08-31 02:41:44 +00:00
* fix: handle caption skin workflow * docs(skills): simplify the finalize step across video workflows - Drop --strict-layout; all skills use plain `hyperframes inspect` - Add the caption text_box_overflow false-positive note to faceless-explainer - On a failed check, the orchestrator makes the cheapest safe edit itself (no worker re-dispatch / Step 3 backtrack language) - Snapshot: glance at the stitched contact-sheet.jpg and move on Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(auth): onboarding-first `auth status` + shared TTS/BGM preflight When no HeyGen credential is configured, `hyperframes auth status` now prints registration-first guidance instead of a terse error: - Interactive / agent-driven sessions get sign-in guidance led by `hyperframes auth login` (the OAuth step that also creates an account and is shared with heygen-cli), and never steer users to a per-repo `.env`. CI / non-interactive runs get a terse note. Exit 1 is kept so the "am I logged in?" `$?` contract still holds. - It probes which local engine voice/music will fall back to (Kokoro / MusicGen, mirroring the skill resolution order) and whether their Python deps are installed, with a pip hint when missing. `--json` exposes `recommended_action` + `offline_engines` for skills to branch. - `doctor` gains matching "TTS (Kokoro)" / "BGM (MusicGen)" checks via the same shared probe (findPython/hasPythonModules extracted to tts/python.ts; provider resolution in audio/providers.ts). Every TTS/BGM workflow now relays this at Step 0 (setup) instead of improvising its own "missing key" prompt: pr-to-video, product-launch- video, faceless-explainer, website-to-video, music-to-video. The canonical behavior + key-priority table live once in hyperframes-media. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(pr-to-video): scale recommended video length to PR change size Step 0 led with a fixed ~60-90s length default. Now the recommended length is derived from the PR's diff stat (lines added+deleted, nudged by file count) on a tier scale (trivial ~20-40s → large ~110-180s, hard cap ~3 min), reusing the same PR peek already done to infer the angle. The agent states the basis when proposing it, and a huge PR with one headline change still stays tight. User can always override. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(captions): embed brand fonts whose files use separators brandFontFaces() matched font files by stripping only whitespace, so an underscore/hyphen-named file (TT_Norms_Pro_Bold.woff2) never matched the family key "ttnormspro" — captions shipped with no @font-face, the font_family_without_font_face bug. Now both family and filename normalize away all non-alphanumerics; families match longest-key-first so a parent family can't swallow a more specific one's files (TT Norms Pro vs Mono); each file is claimed once; "demibold" ranks before "bold"; and when nothing matches it warns loudly at build time instead of returning "". Also: parseFonts() falls back to h1/h2/title/hero display roles, and the frame-worker + caption authoring docs spell out that only shipped font files render — no system CJK/Devanagari families on the headless renderer. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(hyperframes-media): enforce sign-in preflight on standalone BGM/TTS A one-off "generate me a BGM" request went straight to local MusicGen without recommending sign-in: bgm.md/tts.md framed the no-credential path as an automatic fallback, so the generation path bypassed the Preflight stop, and the preflight used a bare `hyperframes auth status` that isn't on PATH in a fresh `npx skills` project. - Preflight now applies to one-off generation as well as workflows, uses `npx hyperframes auth status`, and says: if the CLI can't run, still recommend signing in and STOP — never treat "no credential" as a silent green light for local generation. - bgm.md and tts.md point at the Preflight before generating, reframing local generation as the fallback the user opts into, not a default. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(auth): add Authentication & API keys guide Document signing in, the keys each capability (voice, music, capture) uses, their resolution priority, and the fully local fallback. Add the guide to the nav and cross-link it from the cloud deploy note and the CLI env-var reference. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(lint): strip HTML comments in a fixpoint loop (CodeQL) Single-pass <!-- --> removal can re-form a complete comment from adjacent markers (e.g. `<<!-- -->!-- ... -->`), letting a decoy <template> survive and hijack the template-boundary match. Loop to a fixpoint, mirroring the captions.mjs precedent; add a regression test that fails on single-pass (2 root findings) and passes on the loop. Also wrap the build-frame.mjs node:fs imports to satisfy oxfmt — the new copyFileSync import pushed the line past the width limit, which was the sole cause of the Format / Preflight CI failures. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(lint): strip HTML comments with a linear scan (CodeQL ReDoS) The fixpoint loop still ran a /<!--[\s\S]*?-->/ regex per pass, which backtracks O(n^2) on inputs with many unterminated "<!--" — CodeQL js/polynomial-redos (high). Looping the same regex (the prescribed fix) never addressed this; only the regex itself does. Replace it with an indexOf-based linear strip in utils.ts (stripHtmlComments), kept in a fixpoint loop so markers that re-form when a comment is removed are still stripped. 200k unterminated "<!--" now strips in ~3ms instead of quadratic time; behavior is otherwise unchanged — unterminated comments are kept verbatim, as the old regex left them. The re-forming regression test still guards it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(auth): make TTS/BGM sign-in guidance accurate and runnable From team review of the not-signed-in onboarding: - OAuth is a `hyperframes auth login` feature only. The separate `heygen` CLI is API-key-only — `heygen auth login` stores a pasted key, it is not OAuth and does not create an account. Stop presenting the two CLIs as the same OAuth/sign-up step. - Use `npx hyperframes` in every imperative and runtime hint. Bare `hyperframes` is not on PATH on a fresh machine (command not found); only `npx hyperframes` is guaranteed. Also updates the JSON recommended_action. - Drop `heygen auth login` from the terminal/skill onboarding: it needs its own install and there is no `npx heygen`, so it was a command-not-found trap. The shared-credential fact stays in the reference docs. Covers the `auth status` guidance + tests, the Authentication docs, the shared hyperframes-media preflight (SKILL, requirements, tts, error hints), and the `npx hyperframes auth status` preflight in every TTS/BGM workflow (pr-to-video, product-launch-video, faceless-explainer, website-to-video, music-to-video). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
301 lines
9.5 KiB
TypeScript
301 lines
9.5 KiB
TypeScript
// Shared types, regex constants, and utility functions used across lint rule modules.
|
|
// Nothing in this file should emit findings — it only parses and extracts.
|
|
|
|
export type OpenTag = {
|
|
raw: string;
|
|
name: string;
|
|
attrs: string;
|
|
index: number;
|
|
};
|
|
|
|
export type ExtractedBlock = {
|
|
attrs: string;
|
|
content: string;
|
|
raw: string;
|
|
index: number;
|
|
};
|
|
|
|
const TAG_PATTERN = /<([a-z][\w:-]*)(\s[^<>]*?)?>/gi;
|
|
export const STYLE_BLOCK_PATTERN = /<style\b([^>]*)>([\s\S]*?)<\/style>/gi;
|
|
export const SCRIPT_BLOCK_PATTERN = /<script\b([^>]*)>([\s\S]*?)<\/script>/gi;
|
|
const COMPOSITION_ID_IN_CSS_PATTERN = /\[data-composition-id=["']([^"']+)["']\]/g;
|
|
export const TIMELINE_REGISTRY_INIT_PATTERN =
|
|
/window\.__timelines\s*=\s*window\.__timelines\s*\|\|\s*\{\}|window\.__timelines\s*=\s*\{\}|window\.__timelines\s*\?\?=\s*\{\}/i;
|
|
export const TIMELINE_REGISTRY_ASSIGN_PATTERN =
|
|
/window\.__timelines(?:\[[^\]]+\]|\.[A-Za-z_$][\w$]*)\s*=/i;
|
|
export const WINDOW_TIMELINE_ASSIGN_PATTERN =
|
|
/window\.__timelines(?:\[\s*["']([^"']+)["']\s*\]|\.\s*([A-Za-z_$][\w$]*))\s*=\s*([A-Za-z_$][\w$]*)/i;
|
|
export const INVALID_SCRIPT_CLOSE_PATTERN = /<script[^>]*>[\s\S]*?<\s*\/\s*script(?!>)/i;
|
|
|
|
const TIMELINE_REGISTRY_KEY_PATTERN =
|
|
/window\.__timelines(?:\[\s*["']([^"']+)["']\s*\]|\.\s*([A-Za-z_$][\w$]*))\s*=/g;
|
|
|
|
export function extractOpenTags(source: string): OpenTag[] {
|
|
const tags: OpenTag[] = [];
|
|
let match: RegExpExecArray | null;
|
|
const pattern = new RegExp(TAG_PATTERN.source, TAG_PATTERN.flags);
|
|
while ((match = pattern.exec(source)) !== null) {
|
|
const raw = match[0];
|
|
if (raw.startsWith("</") || raw.startsWith("<!")) continue;
|
|
tags.push({
|
|
raw,
|
|
name: (match[1] || "").toLowerCase(),
|
|
attrs: match[2] || "",
|
|
index: match.index,
|
|
});
|
|
}
|
|
return tags;
|
|
}
|
|
|
|
export function extractBlocks(source: string, pattern: RegExp): ExtractedBlock[] {
|
|
const blocks: ExtractedBlock[] = [];
|
|
let match: RegExpExecArray | null;
|
|
const p = new RegExp(pattern.source, pattern.flags);
|
|
while ((match = p.exec(source)) !== null) {
|
|
blocks.push({
|
|
attrs: match[1] || "",
|
|
content: match[2] || "",
|
|
raw: match[0],
|
|
index: match.index,
|
|
});
|
|
}
|
|
return blocks;
|
|
}
|
|
|
|
/**
|
|
* Find the `<html>` open tag in the source. Distinct from `findRootTag`,
|
|
* which returns the first element inside `<body>` — the latter is "the
|
|
* composition's visible root", whereas `<html>` is where document-level
|
|
* metadata like `data-composition-variables` lives.
|
|
*/
|
|
export function findHtmlTag(source: string): OpenTag | null {
|
|
const match = /<html\b([^<>]*)>/i.exec(source);
|
|
if (!match) return null;
|
|
return {
|
|
raw: match[0],
|
|
name: "html",
|
|
attrs: match[1] ?? "",
|
|
index: match.index,
|
|
};
|
|
}
|
|
|
|
export function findRootTag(source: string): OpenTag | null {
|
|
const bodyOpenMatch = /<body\b([^>]*)>/i.exec(source);
|
|
const bodyCloseMatch = /<\/body>/i.exec(source);
|
|
if (
|
|
bodyOpenMatch &&
|
|
(readAttr(bodyOpenMatch[0], "data-composition-id") ||
|
|
readAttr(bodyOpenMatch[0], "data-width") ||
|
|
readAttr(bodyOpenMatch[0], "data-height"))
|
|
) {
|
|
return {
|
|
raw: bodyOpenMatch[0],
|
|
name: "body",
|
|
attrs: bodyOpenMatch[1] ?? "",
|
|
index: bodyOpenMatch.index,
|
|
};
|
|
}
|
|
const bodyStart = bodyOpenMatch ? bodyOpenMatch.index + bodyOpenMatch[0].length : 0;
|
|
const bodyEnd =
|
|
bodyOpenMatch && bodyCloseMatch && bodyCloseMatch.index > bodyStart
|
|
? bodyCloseMatch.index
|
|
: source.length;
|
|
const bodyContent = bodyOpenMatch ? source.slice(bodyStart, bodyEnd) : source;
|
|
const bodyTags = extractOpenTags(bodyContent);
|
|
for (const tag of bodyTags) {
|
|
if (["script", "style", "meta", "link", "title"].includes(tag.name)) continue;
|
|
return { ...tag, index: tag.index + bodyStart };
|
|
}
|
|
return null;
|
|
}
|
|
|
|
export function readAttr(tagSource: string, attr: string): string | null {
|
|
if (!tagSource) return null;
|
|
const escaped = attr.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
|
|
const match = tagSource.match(new RegExp(`\\b${escaped}\\s*=\\s*["']([^"']+)["']`, "i"));
|
|
return match?.[1] || null;
|
|
}
|
|
|
|
/**
|
|
* Read an attribute that may legitimately contain the opposite quote
|
|
* character. `readAttr` truncates `data-variable-values='{"title":"Hello"}'`
|
|
* at the first internal `"` because its `[^"']+` class excludes both quote
|
|
* types. This variant alternates: a double-quoted value never contains an
|
|
* unescaped `"`, and a single-quoted value never contains an unescaped `'`,
|
|
* so each branch can use a quote-specific class.
|
|
*
|
|
* Use for attributes whose values are JSON or otherwise carry the opposite
|
|
* quote character. Existing single-token attributes (`id`, `class`, etc.)
|
|
* stick with `readAttr` for consistency with the rest of the lint code.
|
|
*/
|
|
export function readJsonAttr(tagSource: string, attr: string): string | null {
|
|
if (!tagSource) return null;
|
|
const escaped = attr.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
|
|
const match = tagSource.match(new RegExp(`\\b${escaped}\\s*=\\s*(?:"([^"]*)"|'([^']*)')`, "i"));
|
|
if (!match) return null;
|
|
return match[1] ?? match[2] ?? null;
|
|
}
|
|
|
|
export function collectCompositionIds(tags: OpenTag[]): Set<string> {
|
|
const ids = new Set<string>();
|
|
for (const tag of tags) {
|
|
const compId = readAttr(tag.raw, "data-composition-id");
|
|
if (compId) ids.add(compId);
|
|
}
|
|
return ids;
|
|
}
|
|
|
|
export function extractCompositionIdsFromCss(css: string): string[] {
|
|
const ids = new Set<string>();
|
|
let match: RegExpExecArray | null;
|
|
const pattern = new RegExp(
|
|
COMPOSITION_ID_IN_CSS_PATTERN.source,
|
|
COMPOSITION_ID_IN_CSS_PATTERN.flags,
|
|
);
|
|
while ((match = pattern.exec(css)) !== null) {
|
|
if (match[1]) ids.add(match[1]);
|
|
}
|
|
return [...ids];
|
|
}
|
|
|
|
export function extractTimelineRegistryKeys(source: string): string[] {
|
|
const keys = new Set<string>();
|
|
let match: RegExpExecArray | null;
|
|
const pattern = new RegExp(
|
|
TIMELINE_REGISTRY_KEY_PATTERN.source,
|
|
TIMELINE_REGISTRY_KEY_PATTERN.flags,
|
|
);
|
|
while ((match = pattern.exec(source)) !== null) {
|
|
const key = match[1] ?? match[2];
|
|
if (key) keys.add(key);
|
|
}
|
|
return [...keys];
|
|
}
|
|
|
|
export function getInlineScriptSyntaxError(source: string): string | null {
|
|
if (!source.trim()) return null;
|
|
try {
|
|
// eslint-disable-next-line no-new-func
|
|
new Function(source);
|
|
return null;
|
|
} catch (error) {
|
|
if (error instanceof Error) return error.message;
|
|
return String(error);
|
|
}
|
|
}
|
|
|
|
// fallow-ignore-next-line complexity
|
|
export function stripJsComments(source: string): string {
|
|
let out = "";
|
|
let i = 0;
|
|
let quote: "'" | '"' | "`" | null = null;
|
|
let escaped = false;
|
|
|
|
while (i < source.length) {
|
|
const ch = source[i] ?? "";
|
|
const next = source[i + 1] ?? "";
|
|
|
|
if (quote) {
|
|
out += ch;
|
|
if (escaped) {
|
|
escaped = false;
|
|
} else if (ch === "\\") {
|
|
escaped = true;
|
|
} else if (ch === quote) {
|
|
quote = null;
|
|
}
|
|
i += 1;
|
|
continue;
|
|
}
|
|
|
|
if (ch === "'" || ch === '"' || ch === "`") {
|
|
quote = ch;
|
|
out += ch;
|
|
i += 1;
|
|
continue;
|
|
}
|
|
|
|
if (ch === "/" && next === "/") {
|
|
out += " ";
|
|
i += 2;
|
|
while (i < source.length && source[i] !== "\n" && source[i] !== "\r") {
|
|
out += " ";
|
|
i += 1;
|
|
}
|
|
continue;
|
|
}
|
|
|
|
if (ch === "/" && next === "*") {
|
|
out += " ";
|
|
i += 2;
|
|
while (i < source.length) {
|
|
const blockCh = source[i] ?? "";
|
|
const blockNext = source[i + 1] ?? "";
|
|
if (blockCh === "*" && blockNext === "/") {
|
|
out += " ";
|
|
i += 2;
|
|
break;
|
|
}
|
|
out += blockCh === "\n" || blockCh === "\r" ? blockCh : " ";
|
|
i += 1;
|
|
}
|
|
continue;
|
|
}
|
|
|
|
out += ch;
|
|
i += 1;
|
|
}
|
|
|
|
return out;
|
|
}
|
|
|
|
// One linear pass that drops every `<!-- … -->` region. Uses indexOf, not a
|
|
// `/<!--[\s\S]*?-->/` regex: that pattern backtracks O(n²) on inputs with many
|
|
// unterminated "<!--" (CodeQL js/polynomial-redos). An unterminated "<!--" with
|
|
// no closing "-->" is kept verbatim, matching the prior regex's no-match behavior.
|
|
function stripHtmlCommentsOnce(source: string): string {
|
|
let out = "";
|
|
let i = 0;
|
|
for (;;) {
|
|
const start = source.indexOf("<!--", i);
|
|
if (start < 0) return out + source.slice(i);
|
|
const end = source.indexOf("-->", start + 4);
|
|
if (end < 0) return out + source.slice(i);
|
|
out += source.slice(i, start);
|
|
i = end + 3;
|
|
}
|
|
}
|
|
|
|
// Strip HTML comments to a fixpoint. A single pass is not enough: deleting one
|
|
// comment can splice adjacent markers into a fresh, complete <!-- … --> (e.g.
|
|
// "<<!-- -->!-- … -->" → "<!-- … -->"), which would otherwise survive and let a
|
|
// commented-out <template>/tag hijack the linter's tag scan.
|
|
export function stripHtmlComments(source: string): string {
|
|
let out = source;
|
|
for (let prev = ""; prev !== out; ) {
|
|
prev = out;
|
|
out = stripHtmlCommentsOnce(out);
|
|
}
|
|
return out;
|
|
}
|
|
|
|
export function extractScriptTextsAndSrcs(scripts: ExtractedBlock[]): {
|
|
texts: string[];
|
|
srcs: string[];
|
|
} {
|
|
const texts = scripts.filter((s) => !/\bsrc\s*=/.test(s.attrs)).map((s) => s.content);
|
|
const srcs = scripts.map((s) => readAttr(`<script ${s.attrs}>`, "src") || "").filter(Boolean);
|
|
return { texts, srcs };
|
|
}
|
|
|
|
export function isMediaTag(tagName: string): boolean {
|
|
return tagName === "video" || tagName === "audio" || tagName === "img";
|
|
}
|
|
|
|
export function truncateSnippet(value: string, maxLength = 220): string | undefined {
|
|
const normalized = value.replace(/\s+/g, " ").trim();
|
|
if (!normalized) return undefined;
|
|
if (normalized.length <= maxLength) return normalized;
|
|
return `${normalized.slice(0, maxLength - 3)}...`;
|
|
}
|