docs(prompting): capstone v4 — dwell-and-sweep rhythm, real brand frame.md, text-behind-subject

This commit is contained in:
Vance Ingalls
2026-07-22 21:11:44 -07:00
parent a90e343e8e
commit b13457be93
+11 -7
View File
@@ -14,7 +14,7 @@ The concept is the product: the viewer travels along one literal timeline — th
Two earlier builds of this capstone passed every gate and still failed review — first as sparse frames floating in space, then as a well-dressed slideshow: discrete scenes, hard cuts, motion confined inside each frame. What fixed it wasn't more animation; it was a **cinematography contract** written into the prompt:
- **One continuous camera.** The whole film is a single wide world the camera dollies across in one unbroken move — easing through regions, accelerating between them, never fully stopping until the final breath. Scene changes happen by *arriving*: the next region is already visible at the frame edge before the camera reaches it, and the previous one exits by parallax, not by fade.
- **One continuous camera, dwell-and-sweep.** The whole film is a single wide world the camera dollies across in one unbroken move — but the rhythm matters as much as the continuity: the camera SWEEPS between regions, then eases to a genuine full stop for a couple of seconds while each region's main point lands (the count completing, the kicker stamping, the spoken keywords), then accelerates away. The dwell is what gives the sweep meaning — without the rest, emphasis dies mid-motion. During a dwell the world keeps resolving; the camera rests, the film never freezes. Scene changes happen by *arriving*: the next region is already visible at the frame edge before the camera reaches it, and the previous one exits by parallax, not by fade.
- **Persistent elements that tell the story.** Three things never leave frame: the **wire** (a single continuous line that is the timeline itself — it becomes the chart baseline, the map arc, the waveform, the 3D coil), the **ruler** (a frame-ruler chrome strip whose playhead tracks the film's real playback time), and the **protagonist** — the `<div class="clip">` typed in the opening, which rides the wire through every region and is the thing that finally renders at the end.
- **Anti-slideshow negatives.** No element cluster that fades in centered, sits, and fades out; no region that reads as a static slide the camera happens to pass; no motion that is just idling in place.
@@ -26,8 +26,10 @@ This is the exact prompt that produced both renders — nothing paraphrased, not
> /general-video Make a ~60-second 1920x1080 promo film for HyperFrames — the open-source framework that turns hand-authored HTML and GSAP animation into deterministic, frame-accurate video. The film's concept IS the product: the viewer travels along one literal timeline — the thing HyperFrames is built on — from the first typed character to the finished MP4. Everything true, nothing invented: real catalog counts, real cloud regions, real product surfaces.
>
> **Design system: the provided `frame.md` is the single source of brand truth — read it first and obey it exactly.** It is the real HyperFrames brand (by HeyGen), dark-first: flat near-black ground (`--page-bg` `#0a0a0a`), greyscale text hierarchy (`#fafafa` / `#a1a1aa` / `#71717a`), and **mint `#3CE6AC` used sparingly as the single accent — mint means go / primary / active**. The mint→cyan gradient (`#3CE6AC → #00E3FF`) appears in exactly one place in the whole film: the final CTA lockup. Typography from the file: **ABC Solar Display 700** for the display lockups ("WRITE HTML." / "WRITE HTML. RENDER VIDEO."), **TT Norms Pro** for everything else, **TT Norms Pro Mono** for all machine-flavored text (code, commands, timecodes, counts, region labels) with `tabular-nums`. Sentence case everywhere except small uppercase letter-spaced eyebrows. Glass surfaces use the file's `--glass-*` tokens with real backdrop blur. Radii, borders, shadows, and easing curves come from the file's tokens (`--ease-out-quint` entrances, `--ease-spring` count-bumps). Font files are provided as woff2 assets — bundle them, don't fetch from the network.
>
> **Cinematography contract (the most important section — this is what the film lives or dies on):**
> - ONE continuous camera journey. The whole film is a single wide world (one large scene space) that the camera dollies across — left to right, with occasional push-ins and one descent — in one unbroken move. The camera is ALWAYS traveling: it eases through regions of interest (slowing to let a beat land, never fully stopping for more than a breath) and accelerates between them. There are NO hard cuts anywhere in the film except inside the two sanctioned seam moments named below.
> - ONE continuous camera journey with a dwell-and-sweep rhythm. The whole film is a single wide world (one large scene space) that the camera dollies across — left to right, with occasional push-ins and one descent — in one unbroken move with NO hard cuts anywhere except inside the two sanctioned seam moments named below. The rhythm: the camera SWEEPS between regions, then eases to a genuine REST at each region's hero moment — a full stop of roughly 1.52.5 seconds while the main point lands (the hero count completing, the kicker stamping, the spoken caption line, the check completing) — then accelerates away. The dwell is what gives the sweep meaning: without the rest, emphasis dies mid-motion. During a dwell the WORLD keeps resolving (counters tick, labels stamp, secondary motion continues) — the camera rests, the film never freezes.
> - Scene changes happen by ARRIVING, not cutting. The next region is already visible at the edge of frame before the camera reaches it; the previous region exits by parallax, not by fade-out.
> - NO slideshow tells: no element cluster that fades in centered, sits, and fades out; no region that reads as a static slide while the camera happens to pass; no motion that is just shaking or idling in place. Motion = the camera's travel + elements entering along the direction of travel + elements the camera hands off between regions.
> - Three persistent elements thread the ENTIRE film and are never absent:
@@ -38,7 +40,7 @@ This is the exact prompt that produced both renders — nothing paraphrased, not
> **Message, stated once:** "Write HTML. Render video." — opening hook and closing lockup. Arc: Type → Mount → Measure → Everywhere → Material → Surface → Rhythm → Depth → Render. Audience: developers evaluating a programmable video tool. Mood: focused build-up that keeps accelerating, resolving confident and warm at the render.
>
> **Direction (every region obeys it):**
> - Two-color discipline: `ground` deep charcoal `#141210`, `ink` vivid amber-orange `#ff7a33` — the only hue; neutral warm off-white for reading text is non-chromatic. Emphasis via inversion, weight, scale, density — never a third color.
> - Brand restraint per `frame.md`: the world is greyscale-on-near-black; `ground` `#0a0a0a` and mint `ink` `#3CE6AC` are declared as the two composition variables. The wire, the playhead, active chips, and each region's single hero accent are mint; everything else is the greyscale text/border hierarchy. No other hue anywhere — except cyan `#00E3FF`, which exists only inside the one sanctioned mint→cyan gradient on the final CTA.
> - VO-paced reveals: each region's elements land on their spoken cues as the camera arrives; secondary elements keep resolving while the camera is present; the region is complete just as the camera accelerates away.
> - Permanent chrome: the ruler (persistent element 2) plus a faint ink grid at ~8% opacity over the ground, everywhere in the world.
> - Density: every region fills three roles — one focal element at display scale, at least two supporting elements on their own cues, and the chrome/wire. Asymmetric compositions; display type ~a tenth of frame width; three depth layers with parallax between them (foreground elements slide past faster than the wire plane, backdrop slower — the parallax is what sells the continuous travel).
@@ -55,7 +57,7 @@ This is the exact prompt that produced both renders — nothing paraphrased, not
>
> **Everywhere (2128s).** The wire arcs and the ground falls away: adapt the **`world-map` registry block** — the dotted world map scrolls beneath while the wire becomes a great-circle route arcing across it. Real render infrastructure, truthfully labeled: pulses light up at AWS and Google Cloud regions (US-EAST-1 · EU-WEST-1 · AP-SOUTH-1) with a "CLOUD RENDER · LAMBDA + CLOUD RUN" kicker — HyperFrames' actual deploy targets. The map tilts away as the wire descends into the next region.
>
> **Material (2839s).** The wire threads into a working canvas where footage is the material on the timeline: a real talking-head clip (generate a short clip of a person speaking one neutral line via the media pipeline's avatar video generation — it is material being processed, not a testimonial; **transcode it to HEVC `hvc1`** so the automatic proxy subsystem carries preview) sits as a clip on the wire. As the camera arrives, the framework visibly works the footage: the subject **mattes out via background removal** (the background peels away as a layer sliding off along the travel direction), a designed lower-third chip docks beside them, and **word-synced captions** from the clip's own transcription land word by word — adapt a bold **catalog caption component** (`caption-kinetic-slam` or similar). The clip's own audio ducks the BGM briefly; the VO resumes as the camera pulls away and the captioned clip recedes by parallax.
> **Material (2839s).** The wire threads into a working canvas where footage is the material on the timeline: a real talking-head clip (generate a short clip of a person speaking one neutral line via the media pipeline's avatar video generation — it is material being processed, not a testimonial; **transcode it to HEVC `hvc1`** so the automatic proxy subsystem carries preview) sits as a clip on the wire. The order of operations IS the story: as the camera arrives and BEFORE the person speaks, the framework mattes the footage — **the background peels away via background removal**, sliding off along the travel direction and leaving the cutout standing alone on the brand ground. THEN they speak, and the main **keywords of their own line — derived from the clip's transcription land word-synced as huge display text BEHIND the cutout**, each keyword slamming in on its spoken moment with the subject's silhouette occluding it (the two-layer text-behind-subject plate); style the keyword type by adapting a bold **catalog caption component** (`caption-kinetic-slam` or similar) at display scale. A designed lower-third chip docks beside them. The clip's own audio ducks the BGM briefly; the VO resumes as the camera pulls away and the worked footage recedes by parallax.
>
> **Surface (3945s).** SANCTIONED SEAM #1: the camera pushes through an **`sdf-iris` shader transition** — the iris opens ONTO the continuation of the same wire (the wire is visible through the iris throughout; this is a lens the journey passes through, not a cut). Inside: frosted-glass inspector panels (real translucency — blur over what's behind; the product's Studio design language) hover above the wire carrying live values (`ease: power3.out`, a color token, an fps readout), over a **generated-artwork backdrop** — a flat abstract mural evoking motion trails, generated **palette-neutral (grayscale) and duotoned live by CSS** from the two variables so it re-skins with everything else. Drifting `data-*` tags ride between mural and glass, clear of all text.
>
@@ -69,13 +71,13 @@ This is the exact prompt that produced both renders — nothing paraphrased, not
>
> **Technique, pinned:** continuous world-camera via GSAP transforms (the dolly is the film); parallax layers in every region; `data-chart`, `world-map`, `sdf-iris`, and a caption component adapted from the registry; per-character typing at the open; matched-motion handoffs at every region boundary (the wire IS the match); full footage pipeline (generated clip → HEVC → auto-proxy → background removal → lower-third → word-synced captions); real beat grid from `hyperframes beats`; real Three.js via the frame adapter; seeded two-frame-hold confetti; SFX on hits; VO-paced reveals throughout.
>
> **Variables:** expose `ground` and `ink` as composition variables on the single root file, bound via CSS custom properties everywhere (including the duotoned mural), so one `--variables` call re-skins the entire journey. It will be rendered twice: the default palette, and a second full render with `{"ground":"#0d1420","ink":"#c8ff3d"}`.
> **Variables:** expose `ground` (default `#0a0a0a`) and `ink` (default `#3CE6AC`) as composition variables on the single root file, bound via CSS custom properties everywhere (including the duotoned mural), so one `--variables` call re-skins the entire journey. It will be rendered twice: the default brand palette, and a second full render with `{"ground":"#0d1420","ink":"#c8ff3d"}`.
>
> **Architecture constraint (technical):** single composition file — one `index.html`, one variable scope. The world is one wide scene space traversed by a camera transform; regions may still be timed `class="clip"` sections synchronized to the camera's arrival (so capture stays efficient), but their positions form one continuous world and the wire + ruler + protagonist chip layers persist across 100% of the film. No `data-composition-src` sub-files.
>
> **Voice & music:** confident, energetic female narration — Kokoro `af_sky`, passed explicitly via `--voice` — riding the journey's momentum (the VO never waits for the camera; the camera arrives on the VO). Driving, beat-forward electronic BGM from the catalog at the engine's default narration-bed volume; the same track feeds the Rhythm region's beat analysis.
>
> **Negatives:** no hard cuts outside the two sanctioned seams; no static regions; no shaking or idle-in-place motion passed off as energy; no third hue; no stock footage or photography; no fictional stats, projects, or testimonials — every number and label must be true of HyperFrames today; no real OS/browser chrome; no unseeded randomness; none of the film is silent.
> **Negatives:** no hard cuts outside the two sanctioned seams; no static regions; no shaking or idle-in-place motion passed off as energy; no hue beyond mint (cyan only inside the one CTA gradient) — and per the brand file, no purple-blue SaaS gradients ever; no stock footage or photography; no fictional stats, projects, or testimonials — every number and label must be true of HyperFrames today; no real OS/browser chrome; no unseeded randomness; none of the film is silent.
## Walking the prompt, chapter by chapter
@@ -87,7 +89,7 @@ Every region pins a different technique, and every technique has a chapter:
| Mount | The timeline's own vocabulary — tracks, clips, `data-*` labels | [Prompt anatomy](/prompting/anatomy) |
| Measure | `data-chart` block; the wire as chart baseline (matched motion) | [Data and maps](/prompting/data-and-maps), [Transitions](/prompting/transitions) |
| Everywhere | `world-map` block; real Lambda / Cloud Run regions | [Data and maps](/prompting/data-and-maps), [Rendering and output](/prompting/rendering-and-output) |
| Material | Generated clip → HEVC proxy → background removal → lower-third → word-synced caption component | [Media and audio](/prompting/media-and-audio), [Caption styles](/prompting/captions-catalog), [Overlays and lower thirds](/prompting/overlays-and-lower-thirds) |
| Material | Generated clip → HEVC proxy → background removed *before* the speech → the line's keywords slam in word-synced as display text BEHIND the cutout (silhouette occlusion) → lower-third | [Media and audio](/prompting/media-and-audio), [Caption styles](/prompting/captions-catalog), [Overlays and lower thirds](/prompting/overlays-and-lower-thirds) |
| Surface | `sdf-iris` shader lens, frosted glass, palette-neutral generated mural | [Transitions](/prompting/transitions), [VFX and liquid glass](/prompting/vfx-and-liquid-glass), [When to generate artwork](/prompting/generated-artwork) |
| Rhythm | Cuts on a real analyzed beat grid (`hyperframes beats`) | [Music videos and slideshows](/prompting/music-and-slideshows) |
| Depth | Real Three.js via the frame adapter, wire as 3D coil | [Runtimes and 3D](/prompting/runtimes-and-3d) |
@@ -95,6 +97,8 @@ Every region pins a different technique, and every technique has a chapter:
The camera itself is [Motion that reads premium](/prompting/motion)'s "camera as actor" rule taken to its logical end — the camera isn't *in* the film, it *is* the film. The persistent wire/ruler/protagonist are [Storyboards](/prompting/storyboards)' callback device stretched across the whole runtime instead of two frames. And the whole thing is written in [High-fidelity looks](/prompting/visual-specs)' density: exact hexes, exact timings, named eases, quoted copy.
One more chapter runs through every pixel: the film's look isn't described in the prompt at all — the prompt points at a **`frame.md` brand file** (the real HyperFrames design system: near-black ground, greyscale hierarchy, mint as the single go/active accent, ABC Solar Display lockups, TT Norms Pro Mono for machine text, one sanctioned mint→cyan gradient on the CTA) and says *obey it exactly*. That's [Design systems and brand](/prompting/design-systems)'s whole thesis — a source of brand truth beats any adjective — demonstrated at film scale.
### Why one file, still
HyperFrames scopes `data-composition-variables` **per composition file**, so the single-file constraint from the earlier capstones carries over — one `index.html`, one variable scope, which is what lets one `--variables` flag re-skin the whole journey. The new wrinkle: the world is one wide scene space traversed by a camera transform, while regions stay timed `class="clip"` sections synchronized to the camera's arrival — the persistent layers (wire, ruler, chip) are the only elements alive for all 62 seconds, so capture stays efficient without breaking the illusion of one continuous space. See [Variables and templating](/prompting/variables-and-templating).