feat: website capture pipeline + 7-step video production skill (#284)

* feat(cli): add website capture with AI-powered DESIGN.md generation

Adds `hyperframes capture <url>` command that extracts a complete design
system from any website, producing AI-agent-ready output:

- Full-page screenshot (lazy-load aware, nav at top)
- AI-generated DESIGN.md via Claude API (colors, typography, elevation,
  components, do's/don'ts) with programmatic asset catalog (136+ assets
  with HTML context annotations like img[src], css url(), link[rel=preload])
- CSS-purged compositions (87% size reduction via PurgeCSS)
- HTML-prettified compositions (one-tag-per-line for AI readability)
- CLAUDE.md + .cursorrules auto-generated for AI agent instructions
- Asset deduplication (srcset variants) and tracking pixel filtering

* feat(cli): add gemini 3.1 pro, playwright screenshots, replica refinement

- switch to gemini 3.1 pro (gemini-3.1-pro-preview) with claude fallback
- playwright for full-page screenshots (fixes puppeteer gradient/fixed bugs)
- replica refinement loop: generate, screenshot, compare, fix
- extract inline svgs (50 max, 10kb each) to assets/svgs/
- extract visible text in dom order for content accuracy
- detect js libraries (gsap, three.js, scrolltrigger) via globals
- improved asset catalog grouping and naming
- reverse-engineered aura system prompt documentation
- comprehensive session handoff doc

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: update session handoff with slack research findings

- key finding: team already wants DESIGN.md integration (James, Bin, Vance)
- skills quality matters enormously - must invoke /hyperframes-compose
- eval infrastructure exists (Abhay's dashboards, Teodora's 78-criteria guide)
- templates at templates/ need study before finalizing skill
- session handoff updated with critical next steps

* refactor(cli): simplify capture pipeline, remove replica generator

* feat(capture): add Lottie detection and WebGL shader extraction

Captures Lottie animations via network interception and WebGL shader
source via gl.shaderSource hooking during site crawl. Updates
website-to-hyperframes skill with asset planning guidance, Lottie/shader
reading instructions, and stronger creative direction for scene planning.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(capture): clean pipeline + shader-first creative workflow

Capture pipeline:
- Remove dead deps (puppeteer-extra, stealth plugin, duplicate devDeps)
- Remove duplicate generateAgentPrompt() call (first lied about DESIGN.md)
- Remove dead canvas-to-image code in htmlExtractor (post canvas removal)
- Parallelize image downloads (batches of 5 via Promise.allSettled)
- Fix pre-existing TS error (match[1] guard in font downloader)
- Default capture output to captures/<hostname>

Skill creative overhaul:
- Add shader transition selection to creative director step (Step 4)
- Add shader wiring instructions to engineer step (Step 5)
- Replace 4-line energy modifiers with visual vocabulary table
- Strip rigid scene-by-scene templates from video-recipes.md
- Strip example fill data from scene plan tables
- Add "read transition refs before planning" instruction
- Add creative ambition language ("how the hell did they make this")

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: add skill architecture redesign spec

Comprehensive redesign of website-to-hyperframes skill and capture
pipeline based on code review findings and Claude Code architecture
research. Key changes: remove AI auto-generation, restructure skill
into phases, embed shader boilerplate in scaffold, fix color format.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: add implementation plan for skill architecture redesign

13-task plan covering: capture pipeline cleanup (remove AI generation,
fix colors to HEX, add asset descriptions, shader-ready scaffold),
skill restructuring (4 phases with artifact gates), and compose skill
Visual Identity Gate upgrade.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor(capture): remove AI auto-generation and SDK dependencies

* fix(capture): convert extracted colors to HEX format

* refactor(capture): remove AI key path, add asset descriptions generator

* refactor(capture): update agent prompt, remove hasDesignMd, add asset descriptions

* feat(capture): pre-wire shader transitions in index.html scaffold

* chore: remove duplicate visual-styles.md (canonical is in hyperframes/)

* refactor(skill): rewrite website-to-hyperframes as phase-based orchestrator

* feat(skill): add Phase 1 understand reference

* feat(skill): add Phase 2 design reference with full DESIGN.md schema

* feat(skill): add Phase 3 creative direction reference

* feat(skill): add Phase 4 build reference with inline shader example

* feat(skill): upgrade Visual Identity Gate to produce full DESIGN.md

* docs: update CLAUDE.md skill references for phase-based workflow

* fix: address code review findings

- Remove orphaned `false` argument in generateAgentPrompt call (critical:
  was shifting hasLottie, hasShaders, catalogedAssets parameters)
- Add HSL color handling in rgbToHex via temp element resolution
- Remove build artifact commit section from phase-4-build.md
- Fix __GSAP_TIMELINE reference to __timelines

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(capture): regex double-escape + simplify scaffold + fix asset descriptions

- Double-escape regex in tokenExtractor template literal (\s→\\s, \d→\\d, \(→\\()
  so browser receives valid regex patterns via page.evaluate()
- Simplify index.html scaffold: scene slots + audio + timeline + comment pointing
  to shader-setup.md reference (no broken inline shader boilerplate)
- Fix asset descriptions: use CatalogedAsset.contexts/notes instead of
  nonexistent htmlContext field

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: code review — 16 bugs, 7-step skill rewrite, cleanup

Code fixes:
- snapshot.ts: path traversal guard, browser leak (try/finally), div-by-zero
  for --frames 1, port bind error handling, rAF-based render settle
- index.ts: remove invalid thinkingConfig for gemini-2.5-flash, fix Gemini
  batch/rate-limit comments, fix video preview viewport y-coordinate
- tokenExtractor.ts: remove dead seen[si] dedup code
- gsap.ts: index ALL classes for inline-style transform conflict detection

Skill architecture rewrite (4-phase → 7-step):
- Replace phase-1 through phase-4 with step-1 through step-7
- Add techniques.md (10 visual techniques with code patterns)
- Fix /hyperframes-compose → /hyperframes (skill doesn't exist)
- Fix captures/arc-browser reference → shader-setup.md (file doesn't exist)
- Fix step-7 hardcoded captures/stripe path
- Document Gemini API free/paid rate limits in step-1

Cleanup:
- CLAUDE.md: restore from Stripe-capture overwrite, update 4-phase → 7-step
- .gitignore: add PR #267 skills (hyperframes-animation-map, hyperframes-contrast)
- Delete old phase-*.md, animation-recreation.md, tts-integration.md

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: remove dev artifacts, research docs, wrong lockfiles

Remove files that shouldn't ship in this PR:
- docs/research/ (aura analysis, prompt catalogs)
- docs/session-*.md, docs/SESSION-HANDOFF.md (dev notes)
- docs/superpowers/ planning and spec docs
- pnpm-lock.yaml at root and cli (repo uses bun, not pnpm)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(CLAUDE.md): align with main — slim format, add website-to-hyperframes mention

Main PR #283 removed the full skills table from CLAUDE.md and moved it
to AGENTS.md. Align with that decision: use main's slim dev-focused
format, fix pnpm→bun references, add one-line /website-to-hyperframes
pointer.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(cli): add capture command to help groups

The capture command was registered in cli.ts but missing from
the help groups, so it wouldn't appear in `hyperframes --help`.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: format skill reference files (oxfmt)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: regenerate bun.lock after rebase

The lockfile was stale after rebasing onto main — bun install
--frozen-lockfile failed in CI because new dependencies (google/genai,
patchright, purgecss) weren't reflected in the lockfile.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR review comments + improve capture quality

Review fixes (16 comments from jrusso1020 + vanceingalls):
- screenshotCapture: remove Playwright dep, use Puppeteer for all screenshots
- screenshotCapture: dynamic screenshot count based on page height (30% overlap)
- snapshot.ts: fix duration() function-vs-property bug, cross-platform path guard
- htmlExtractor: fix code injection via parameterized evaluate
- index.ts: video preview re-measures position after scroll, .env file loading
- capture.ts: BLOCKED.md on timeout failures
- gsap.ts: 5 inline-style lint tests added (all pass)
- Remove Playwright, patchright deps; @google/genai to optionalDependencies
- Gitignore: generic patterns instead of 20 hardcoded directories
- Remove asset-sourcing.md, video-recipes.md (unused, duplicated guidance)

Capture quality improvements (tested on 10+ websites):
- Color extraction: canvas-based oklch/lab resolver, pixel sampling via
  elementFromPoint, broad sweep for accent colors, gradient/shadow extraction
- Section detection: broadened selectors for div-based layouts, height cap
  to skip page-level wrappers, parent bg walkup for dark sites
- Font downloads: cap 6 per family / 30 total (Cal.com: 306→30)
- CTA detection: text pattern matching + nav context filtering
- Heading text: innerText with whitespace normalization
- Gemini captioning: maxOutputTokens 100→300, .env auto-loading
- .env.example updated with GEMINI_API_KEY docs
- TTS ranking: Kokoro first with Python 3.10+ note

* fix: address PR review comments + improve capture quality

Review round 2 fixes (jrusso1020 + vanceingalls):
- verify/index.ts: add path traversal guard (relative + isAbsolute)
- verify/index.ts: fix sections[i] undefined typecheck error (CI green)
- index.ts: escape Lottie JSON with \u003c to prevent </script> breakout
- step-4-storyboard: fix technique count contradiction (2-3 per beat, not
  across whole video)
- step-6-build: perspective tilt uses gsap.set() instead of CSS transform
  (avoids GSAP overwrite conflict)
- step-1-capture: reorder — command first, Gemini note after (zero-config
  is the default path, API key is optional enhancement)
- step-7-validate: add tsx fallback for snapshot command
- step-3-script: vary hook patterns, don't default to number every time
- assetDownloader: exempt SVGs from 10KB minimum filter (company logos
  like Hubspot/Intel/DHL are 2-6KB; HeyGen capture: 13→75 assets)

Note: adm-zip was NOT removed (reviewer #3) — it's still in
packages/cli/package.json:30. The root package.json had patchright
and purgecss removed, not adm-zip.

Note: ANTHROPIC_API_KEY not restored in .env.example — grep confirms
zero references in the entire codebase. The @anthropic-ai/sdk dependency
was removed earlier in this branch.

* refactor(capture): split index.ts (1175 to 566 lines) into modules

Mechanical extraction, zero logic changes.

New files:
- mediaCapture.ts (345 lines): Lottie preview, video manifest/screenshots
- contentExtractor.ts (314 lines): library detection, text, Gemini, asset descriptions
- scaffolding.ts (135 lines): .env loading, project scaffold generation

Also fixes false-positive BLOCKED.md with structural Cloudflare detection.
Tested on 20 websites, pre/post output identical.

* chore(capture): remove --split flow (splitter, verify, cssPurger, purgecss)

The --split feature auto-generates compositions from captured HTML — a
different approach from the /website-to-hyperframes skill workflow where
agents build compositions from scratch using the storyboard.

No skill file, no step reference, and no test session ever used --split.
Removes 923 lines of unused code + purgecss dependency.

Backed up to ~/Desktop/capture-split-backup/ for reference.

* fix(security): add ssrf protection, lottie injection fix, oom guard

- assetDownloader: add isPrivateUrl() guard blocking private IP ranges
  (127.x, 10.x, 172.16-31.x, 192.168.x, 169.254.x), cloud metadata
  endpoints, localhost, and non-HTTP schemes
- mediaCapture: fix Lottie JSON injection by loading shell HTML first
  then passing animation data via parameterized page.evaluate()
- index.ts: check Content-Length header before response.buffer() in
  Lottie network interception to avoid OOM on multi-GB responses

* fix(capture): security fixes, timeout, sub-agent dispatch instructions

Security (from miguel-heygen review):
- assetDownloader: export isPrivateUrl() SSRF guard
- htmlExtractor: add isPrivateUrl check before CSS fetch
- mediaCapture: add isPrivateUrl check before Lottie fetch
- mediaCapture: fix previewPage leak (try/finally)
- mediaCapture: skip Lottie files > 2MB for preview (CDP limit)
- contentExtractor: skip images > 4MB for Gemini captioning
- index.ts: check Content-Length before response.buffer() (OOM guard)
- snapshot.ts: register error handler before server.listen()

Capture improvements:
- Default timeout 30s to 120s (Shopify needs ~90s for Cloudflare)
- step-6-build: sub-agent dispatch template with explicit rules:
  pass file PATHS not contents, use local fonts not Google Fonts,
  verify ../assets/ references after each beat

* fix(capture): catalog before DOM mutation, networkidle2, faster Gemini

Critical: asset cataloger now runs BEFORE extractHtml which converts img
src to data URLs. Framer sites like heykuba.com went from 2 to 78 images.

- networkidle2 instead of networkidle0 (unblocks SPAs with WebSockets)
- Lazy-load wait: scroll to bottom, wait for img.complete
- CSS background-image cataloging for Framer/Webflow
- SVG naming: checks class, id, parent, inner text (not just aria-label)
- Gemini batch 5->20, pause 12s->2s (paid tier: 2000 RPM, ~0.001/img)
- maxOutputTokens 300->500, descriptions sorted captioned-first
- Remove tsx fallback from step-1 (reviewer nit, published CLI has it)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
Ular Kimsanov
2026-04-16 12:03:16 -07:00
committed by GitHub
co-authored by Claude Opus 4.6
parent ebc12f7dc9
commit 87f4c77e2f
35 changed files with 5431 additions and 16 deletions
@@ -0,0 +1,316 @@
/**
* Content extraction helpers for the website capture pipeline.
*
* Handles library detection, visible text extraction, Gemini captioning,
* and asset description generation.
*
* All page.evaluate() calls use string expressions to avoid
* tsx/esbuild __name injection (see esbuild issue #1031).
*/
import type { Page } from "puppeteer-core";
import { readdirSync, statSync, readFileSync } from "node:fs";
import { join } from "node:path";
import type { CatalogedAsset } from "./assetCataloger.js";
import type { DesignTokens } from "./types.js";
/**
* Detect JS libraries via window globals, DOM fingerprints, script URLs,
* and WebGL shader analysis.
*
* Returns a deduplicated list of detected library names.
*/
export async function detectLibraries(
page: Page,
capturedShaders?: Array<{ type: string; source: string }>,
): Promise<string[]> {
let detectedLibraries: string[] = [];
try {
detectedLibraries = (await page.evaluate(`(() => {
var libs = [];
function add(name) { if (libs.indexOf(name) === -1) libs.push(name); }
// 1. Window globals (works for CDN-loaded / non-bundled libraries)
if (typeof window.gsap !== 'undefined' || typeof window.TweenMax !== 'undefined') add('GSAP');
if (typeof window.ScrollTrigger !== 'undefined') add('GSAP ScrollTrigger');
if (typeof window.THREE !== 'undefined') add('Three.js');
if (typeof window.PIXI !== 'undefined') add('PixiJS');
if (typeof window.BABYLON !== 'undefined') add('Babylon.js');
if (typeof window.Lottie !== 'undefined' || typeof window.lottie !== 'undefined') add('Lottie');
if (typeof window.__NEXT_DATA__ !== 'undefined') add('Next.js');
if (typeof window.__NUXT__ !== 'undefined') add('Nuxt');
if (typeof window.Webflow !== 'undefined') add('Webflow');
// 2. DOM fingerprints (survive bundling — most reliable for modern sites)
// Three.js sets data-engine on every canvas it creates
var threeCanvas = document.querySelector('canvas[data-engine*="three"]');
if (threeCanvas) add('Three.js (' + (threeCanvas.getAttribute('data-engine') || '') + ')');
// Babylon.js also sets data-engine
var babylonCanvas = document.querySelector('canvas[data-engine*="Babylon"]');
if (babylonCanvas) add('Babylon.js');
// Lottie web components
if (document.querySelector('dotlottie-wc, lottie-player, dotlottie-player')) add('Lottie');
// Rive
if (document.querySelector('canvas[class*="rive"], rive-canvas')) add('Rive');
// React/Next.js
if (document.getElementById('__next')) add('Next.js');
if (document.getElementById('__nuxt')) add('Nuxt');
if (document.querySelector('[data-reactroot], [data-react-helmet]')) add('React');
// Svelte
if (document.querySelector('[class*="svelte-"]')) add('Svelte');
// Tailwind (utility class detection)
if (document.querySelector('[class*="flex "], [class*="grid "], [class*="px-"], [class*="py-"]')) add('Tailwind CSS');
// Framer Motion
if (document.querySelector('[style*="--framer-"], [data-framer-component-type]')) add('Framer Motion');
// 3. Script URL patterns
document.querySelectorAll('script[src]').forEach(function(s) {
var src = s.src.toLowerCase();
if (src.includes('gsap') || src.includes('tweenmax') || src.includes('greensock')) add('GSAP');
if (src.includes('scrolltrigger')) add('GSAP ScrollTrigger');
if (src.includes('three.module') || src.includes('three.min')) add('Three.js');
if (src.includes('pixi')) add('PixiJS');
if (src.includes('lottie') || src.includes('bodymovin')) add('Lottie');
if (src.includes('framer-motion')) add('Framer Motion');
if (src.includes('anime.min') || src.includes('animejs')) add('Anime.js');
if (src.includes('matter.min') || src.includes('matter-js')) add('Matter.js');
if (src.includes('lenis')) add('Lenis (smooth scroll)');
});
return libs;
})()`)) as string[];
} catch {
// Non-blocking
}
// 4. Shader fingerprinting — infer WebGL framework from captured GLSL
try {
const shaders = capturedShaders || [];
if (shaders.length > 0) {
const allSource = shaders.map((s) => s.source).join("\n");
const add = (name: string) => {
if (!detectedLibraries.includes(name)) detectedLibraries.push(name);
};
add("WebGL");
// Three.js shader fingerprints (built-in uniforms that survive bundling)
if (allSource.includes("modelViewMatrix") && allSource.includes("projectionMatrix"))
add("Three.js (confirmed via shaders)");
// PixiJS shader fingerprints
else if (
allSource.includes("vTextureCoord") &&
allSource.includes("uSampler") &&
!allSource.includes("modelViewMatrix")
)
add("PixiJS (confirmed via shaders)");
// Babylon.js shader fingerprints
else if (allSource.includes("viewProjection") && allSource.includes("world"))
add("Babylon.js (confirmed via shaders)");
}
} catch {
/* non-blocking */
}
return detectedLibraries;
}
/**
* Extract all visible text from the page in DOM order using a TreeWalker.
* Truncates to ~30K chars to avoid blowing up downstream prompts.
*/
export async function extractVisibleText(page: Page): Promise<string> {
let visibleTextContent = "";
try {
visibleTextContent = (await page.evaluate(`(() => {
var walker = document.createTreeWalker(document.body, NodeFilter.SHOW_TEXT, null);
var texts = [];
var node;
while (node = walker.nextNode()) {
var text = (node.textContent || '').trim();
if (text.length < 3) continue;
var el = node.parentElement;
if (!el) continue;
var style = getComputedStyle(el);
if (style.display === 'none' || style.visibility === 'hidden' || style.opacity === '0') continue;
var tag = el.tagName.toLowerCase();
if (tag === 'script' || tag === 'style' || tag === 'noscript') continue;
texts.push(text);
}
return texts.join('\\n');
})()`)) as string;
// Truncate to ~30K chars to avoid blowing up the prompt
if (visibleTextContent.length > 30000) {
visibleTextContent = visibleTextContent.slice(0, 30000) + "\n[...truncated]";
}
} catch {
// Non-blocking
}
return visibleTextContent;
}
/**
* Caption downloaded images using Gemini vision API.
*
* Batches requests to stay under free-tier rate limits.
* Returns a map of filename -> caption string.
*/
export async function captionImagesWithGemini(
outputDir: string,
progress: (stage: string, detail?: string) => void,
warnings: string[],
): Promise<Record<string, string>> {
const geminiCaptions: Record<string, string> = {};
const geminiKey = process.env.GEMINI_API_KEY || process.env.GOOGLE_API_KEY;
if (!geminiKey) return geminiCaptions;
progress("design", "Captioning images with Gemini vision...");
try {
const { GoogleGenAI } = await import("@google/genai");
const ai = new GoogleGenAI({ apiKey: geminiKey });
const imageFiles = readdirSync(join(outputDir, "assets")).filter((f: string) =>
/\.(png|jpg|jpeg|webp|gif)$/i.test(f),
);
// Caption in parallel batches via Gemini vision API.
// Free tier: 5 RPM → batch 5, 12s pause (~$0 but slow)
// Paid tier: 2000 RPM → batch 20, 1s pause (~$0.001/image, fast)
// We try a larger batch first; if rate-limited, fall back to smaller batches.
const model = "gemini-2.5-flash";
const BATCH_SIZE = 20;
for (let i = 0; i < imageFiles.length; i += BATCH_SIZE) {
const batch = imageFiles.slice(i, i + BATCH_SIZE);
const results = await Promise.allSettled(
batch.map(async (file: string) => {
const filePath = join(outputDir, "assets", file);
const stat = statSync(filePath);
if (stat.size > 4_000_000) return { file, caption: "" }; // skip images > 4 MB (Gemini inline limit)
const buffer = readFileSync(filePath);
const base64 = buffer.toString("base64");
const ext = file.split(".").pop()?.toLowerCase() || "png";
const mimeType = ext === "jpg" ? "image/jpeg" : `image/${ext}`;
const response = await ai.models.generateContent({
model,
contents: [
{
role: "user",
parts: [
{ inlineData: { mimeType, data: base64 } },
{
text: "Describe this website image in ONE short sentence for a video storyboard. Focus on: what it shows, dominant colors, whether background is light or dark. Be factual, not creative.",
},
],
},
],
config: { maxOutputTokens: 500 },
});
return { file, caption: response.text?.trim() || "" };
}),
);
for (const result of results) {
if (result.status === "fulfilled" && result.value.caption) {
geminiCaptions[result.value.file] = result.value.caption;
}
}
// Pace requests to stay under free tier rate limits (5 RPM for gemini-2.5-flash)
if (i + BATCH_SIZE < imageFiles.length) {
await new Promise((r) => setTimeout(r, 2000)); // 2s pause between batches — paid tier handles 2000 RPM, free tier retries via Promise.allSettled
}
progress(
"design",
`Captioned ${Math.min(i + BATCH_SIZE, imageFiles.length)}/${imageFiles.length} images...`,
);
}
progress("design", `${Object.keys(geminiCaptions).length} images captioned with Gemini`);
} catch (err) {
warnings.push(`Gemini captioning failed: ${err}`);
}
return geminiCaptions;
}
/**
* Generate asset-descriptions.md — one-line descriptions for each downloaded asset.
*
* Returns the description lines (without the markdown header).
*/
export function generateAssetDescriptions(
outputDir: string,
tokens: DesignTokens,
catalogedAssets: CatalogedAsset[],
geminiCaptions: Record<string, string>,
): string[] {
// Sort: Gemini-captioned images first (richest descriptions), then uncaptioned, then SVGs, then fonts
const captionedLines: string[] = [];
const uncaptionedLines: string[] = [];
const svgLines: string[] = [];
const fontLines: string[] = [];
// Describe downloaded images
const assetsPath = join(outputDir, "assets");
try {
for (const file of readdirSync(assetsPath)) {
if (file === "svgs" || file === "fonts" || file === "lottie" || file === "videos") continue;
const filePath = join(assetsPath, file);
const stat = statSync(filePath);
if (!stat.isFile()) continue;
const sizeKb = Math.round(stat.size / 1024);
const catalogMatch = catalogedAssets.find(
(a) => a.url && file.includes(a.url.split("/").pop()?.split("?")[0]?.slice(0, 20) || "___"),
);
const desc = catalogMatch?.description || catalogMatch?.notes || "";
const heading = catalogMatch?.nearestHeading || "";
const section = catalogMatch?.sectionClasses || "";
const aboveFold = catalogMatch?.aboveFold ? "above fold" : "";
const geminiCaption = geminiCaptions[file];
const cleanName = file.replace(/\.[^.]+$/, "").replace(/[-_]/g, " ");
const parts = [`${file}${sizeKb}KB`];
if (geminiCaption) {
parts.push(geminiCaption);
captionedLines.push(parts.join(", "));
} else {
if (desc) parts.push(`"${desc.slice(0, 80)}"`);
if (heading) parts.push(`section: "${heading.slice(0, 60)}"`);
else if (section) parts.push(`in: ${section.split(" ").slice(0, 3).join(" ")}`);
if (aboveFold) parts.push(aboveFold);
if (!desc && !heading) parts.push(cleanName);
uncaptionedLines.push(parts.join(", "));
}
}
} catch {
/* no assets dir */
}
// Describe SVGs
try {
const svgsPath = join(assetsPath, "svgs");
for (const file of readdirSync(svgsPath)) {
if (!file.endsWith(".svg")) continue;
const svgMatch = tokens.svgs.find(
(s) =>
s.label &&
file.includes(
s.label
.toLowerCase()
.replace(/[^a-z0-9]/g, "-")
.slice(0, 15),
),
);
const label = svgMatch?.label || file.replace(".svg", "").replace(/-/g, " ");
const isLogo = svgMatch?.isLogo || file.includes("logo");
svgLines.push(`svgs/${file}${isLogo ? "logo: " : "icon: "}${label}`);
}
} catch {
/* no svgs dir */
}
// Describe fonts
try {
const fontsPath = join(assetsPath, "fonts");
for (const file of readdirSync(fontsPath)) {
fontLines.push(`fonts/${file} — font file`);
}
} catch {
/* no fonts dir */
}
return [...captionedLines, ...uncaptionedLines, ...svgLines, ...fontLines];
}