Miguel R4 blocker on #2359: my R3 fix at renderOrchestrator only updated
the observability copy, leaving the authoritative captureForceScreenshot
local at compileResult.forceScreenshot (false for auto→software). The
frameCapture side clamped its own local and correctly routed screenshot,
but downstream orchestrator code overwrote observability back to
beginframe from the still-false local at two sites:
- Parallel-stream label at renderOrchestrator.ts:2293 mis-labelled the
stream as 'beginframe' when actual capture was 'screenshot'.
- capture_strategy telemetry at renderOrchestrator.ts:2440-2450
overwrote the earlier observability correction, so the final
captureMode observation flipped back to 'beginframe' while the
engine actually captured screenshot.
Fix: extract the clamp into a caller-facing helper
applyConcreteGpuScreenshotClamp(current, resolvedGpuMode, cfg) that
returns the (possibly-promoted) new boolean. Callers assign it back to
their authoritative local, so routing + telemetry + strategy code read
one value.
Changes:
- packages/engine/src/config.ts: new exported
applyConcreteGpuScreenshotClamp; delegates to
shouldClampToScreenshotForConcreteGpu but computes the caller's
final value, not just the clamp decision. Reads the programmatic
opt-out from cfg.forceScreenshotExplicitlyOptedOut. Idempotent on
already-true input.
- packages/engine/src/index.ts: export the new helper.
- packages/engine/src/services/frameCapture.ts: replace the inline
OR expression with applyConcreteGpuScreenshotClamp.
- packages/producer/src/services/renderOrchestrator.ts: assign result
into the AUTHORITATIVE captureForceScreenshot local (was updating
only observability). Downstream parallel-stream label at :2293 and
capture_strategy telemetry at :2440-2450 now read the corrected
value.
Tests: 6 new caller-level cases for applyConcreteGpuScreenshotClamp
covering the exact matrix Miguel called out:
- resolved software + default false → promotes to true (screenshot)
- resolved software + programmatic opt-out → stays false (BeginFrame)
- resolved hardware + default false → stays false
- resolved software + already-true → stays true (idempotent)
- resolved software + env PRODUCER_FORCE_SCREENSHOT=false → stays false
- resolved software + undefined cfg → promotes to true (frameCapture path)
Local: 67/67 engine config tests pass (was 61). oxfmt clean.
Miguel R3 blocker on #2359: the runtime helper only checked the env opt-out
(PRODUCER_FORCE_SCREENSHOT=false), silently defeating the documented
programmatic escape hatch (overrides.forceScreenshot === false) on the
browserGpuMode:'auto' → software probe path. At the concrete-resolution
site the boolean forceScreenshot === false is ambiguous between default
and explicit opt-out — resolveConfig sees the provenance but the runtime
helper does not.
Fix: persist provenance on the resolved config.
- New INTERNAL EngineConfig field forceScreenshotExplicitlyOptedOut, set
by resolveConfig when EITHER env or programmatic explicit-false is
present. Purpose-documented in the type as 'not intended to be set by
callers'.
- shouldClampToScreenshotForConcreteGpu gains an opts.programmaticOptOut
parameter; returns false early when set. Env stays as the third arg
(backward compatibility with existing tests).
- frameCapture.ts and renderOrchestrator.ts pass
config.forceScreenshotExplicitlyOptedOut through at both call sites, so
the auto→software probe path preserves the same escape hatches as
literal browserGpuMode:'software'.
New tests: 5 additional cases across the helper (programmatic opt-out
alone; programmatic beats missing env) and resolveConfig provenance
(programmatic sets flag; env sets flag; neither leaves it undefined).
Local: 61/61 engine config tests pass (was 56).
Matches the isPathInside pattern in fileServer.ts. The helper is called from
webmAlphaCheck.ts's own webmAlphaAdvisory (same file), but fallow doesn't
count intra-file consumption or test-file imports. Suppress the false
positive rather than dropping the direct unit test — the 1/2/3-frame
byte-count gate + per-frame stride logic is load-bearing enough to warrant
independent tests, not just coverage through the outer advisory function.
Addresses Miguel's R1 blockers:
1. `browserGpuMode: "auto"` that runtime-probes to software slipped past the
`resolveConfig` clamp — that clamp only sees the pre-resolve string. Add
`shouldClampToScreenshotForConcreteGpu(resolvedGpuMode, currentForceScreenshot, env)`
in `packages/engine/src/config.ts` and apply it at BOTH concrete-resolution
sites:
- `packages/engine/src/services/frameCapture.ts`: downgrades `preMode`
from "beginframe" to "screenshot" when resolved GPU is software (respects
`PRODUCER_FORCE_SCREENSHOT=false` env opt-out), fixing the routing.
- `packages/producer/src/services/renderOrchestrator.ts`: updates
`captureObservability.forceScreenshot` (and thus `captureMode`) at the
same call site, fixing the observability truth on the auto → software
case.
2. New unit tests in `config.test.ts`:
- Documents the auto-branch gap (resolveConfig leaves auto as
forceScreenshot=false — the runtime companion closes it).
- 5 branch tests on `shouldClampToScreenshotForConcreteGpu` covering
software / hardware / already-forced / env-opt-out / non-"false" env
values.
Full suite: 56/56 pass.
Scope narrowing on Blocker 2: the distributed rendering path at
`packages/producer/src/services/distributed/plan.ts:753-754` and
`renderChunk.ts:462-466` explicitly hardcodes `browserGpuMode:"software",
forceScreenshot:false` post-resolveConfig and stays outside this PR's
invariant boundary. `compileStage` may still flip it to true for alpha
formats, but generic MP4 distributed renders on SwiftShader hosts remain
BeginFrame. That's a separate architectural cleanup (needs its own
behavior-change trace); the PR body now scopes the invariant to the
in-process CLI/orchestrator path.
Address Miguel's R1 blocker + Rames/Miga's testability nit:
- The `-frames:v 3` decode samples AT MOST 3 frames; a legitimate 1- or
2-frame WebM (256 or 512 bytes) was returned as `undefined` (probe
failure), silently skipping the advisory even when every available
pixel was opaque. Accept any positive whole-frame byte count ≤ 768
(multiples of 256), distinguishing successful short-EOF from partial/
malformed decode.
- Export `sampledAlphaIsFullyOpaque` and add 11 direct tests covering:
3/2/1-frame opaque decodes → true; transparent pixel at pos 0 or
final byte → false (guards the alpha-byte stride); non-frame-multiple
/ over-3-frame / zero byte counts → undefined; execFileSync throw →
undefined; findFFmpeg missing → undefined; and one args-shape guard
pinning the load-bearing `-c:v libvpx-vp9` before `-i` (without which
the default decoder silently discards VP9 alpha and the whole check
would false-positive on genuinely-transparent WebMs).
- Update advisory wording from "3 sampled decoded frames" to "every
sampled decoded pixel" so the message is honest for short WebMs.
18/18 tests pass locally under `vitest run`.
Review feedback on #2358: the batch-miss RENDER_FAILED in
runAssetImportMany (asset.ts) throws the same typed error as
client.ts's single-node renderNode, but wasn't labeled — so
cli_error.endpoint would silently come back undefined for the
flow that most heavily exercises /v1/images.
cli_error had no way to tell which figma REST call (images, files_nodes,
variables_local, styles, ...) actually hit RATE_LIMITED/FORBIDDEN/etc, so
the dashboard could see failures spike but not which call caused them.
FigmaClientError now carries a low-cardinality endpoint label (never the
raw fileKey/nodeId), threaded through to cli_error's endpoint property.
Extends webmAlphaCheck.ts (from #2044) with a pixel-level decode probe.
After the tag check passes, decodes 3 sampled frames via
`ffmpeg -c:v libvpx-vp9 -pix_fmt rgba -f rawvideo` at 8x8 and emits a
distinct advisory if every alpha byte reads 255.
#2044 detects the "tag absent" failure mode (ffprobe shows no
`alpha_mode` in stream tags). It doesn't catch a stricter case reported
on CLI 0.7.56 / Windows 11: ALPHA_MODE=1 present but BlockAdditional
alpha side data empty. Under current logic webmAlphaAdvisory sees
`alphaMode: true` and stays silent, so the render ships as opaque
without any signal.
The -metadata:s:v:0 alpha_mode=1 push is a muxer directive that some
ffmpeg builds write unconditionally, independent of whether libvpx-vp9
emitted the alpha plane. Tag presence is necessary but not sufficient
evidence of preserved alpha.
Advisory text names both possibilities (opaque composition OR silent
alpha drop) plus the concrete workaround (png-sequence + prores repack).
Fast path (no tag or missing tag) is unchanged. Probe adds ~1s per
WebM render only when the tag says alpha.
When `browserGpuMode === "software"`, set `forceScreenshot = true` in
`resolveConfig`. Explicit opt-outs (`PRODUCER_FORCE_SCREENSHOT=false`
or `overrides.forceScreenshot === false`) are honored.
This is defense-in-depth on top of the existing platform gates:
1. Linux + software (SwiftShader host) skips BeginFrame, avoiding the
compositor stall on shader-heavy frames under CPU raster (same
motivation as the closed PR #822).
2. `renderOrchestrator`'s reported `captureMode` field is derived from
`cfg.forceScreenshot ? "screenshot" : "beginframe"` — without this
clamp it misreports `"beginframe"` for the actual screenshot capture
on darwin + software.
3. Any new BeginFrame or drawElement entry point that forgets to gate
on GPU mode still routes to screenshot here.
Does NOT fix SwiftShader-on-darwin text-rasterization artifacts (an
ANGLE-SwiftShader issue on macOS text — the fix there is to use
`--browser-gpu`, which routes to `--use-angle=metal`).
Adopt the fake-timer pattern the sibling "recovers when a crashed reclaimer
leaves both lock directories" test in the same file already uses. Without
fake timers, if the dynamic `import("./manager.js")` beat between
`installFsMocks({ initialMtimeMs: Date.now() })` and the `withInstallLock`
call exceeds `staleMs` (50 ms on a busy shared runner with `vi.resetModules()`
per beforeEach), the new immediate-stale short-circuit added in #2328 fires
on iteration 1, breaks out before `waitedMs` reaches `waitNoticeMs=20`, and
no "Waiting for another hyperframes process" warn ever emits — the assertion
at `manager.test.ts:428` (`expected false to be true`) then fails.
Under fake timers, `Date.now()` is frozen at the mtime seed, so the lock
stays non-stale across the dynamic import and the wait-notice branch
observes real polling; `vi.advanceTimersByTimeAsync(staleMs + pollMs * 5)`
then drives the loop past both the wait-notice threshold and the stale
deadline so the reclaim + acquisition still resolves.
Test-only change; no production-code diff. Verified 27/27 in
`packages/cli/src/browser/manager.test.ts` under `vitest run`.
The `hyperframes feedback` convention only prompted for a free-text
`--comment` "with the failing composition pattern and what you tried".
Agents dutifully filed vague reports (blank CJK text, mid-run exit, 4K
timeout) with no error string, no failure-mode, and — critically — no
published composition, so none could be reproduced or root-caused.
Two additions to the CLI skill:
- Lead bug reports with `--file-issue` (+ `--dir`), which publishes a
minimal repro of the project to a public URL. A comment alone almost
never lets a maintainer reproduce; the composition is what does.
- Give the `--comment` a concrete bug checklist: exact error string
verbatim + whether output was produced / fell back / hard-exited; the
isolated trigger; exact command + HF_*/PRODUCER_* env; frame/timestamp +
visual defect. Drop the "repeat env" ask (the CLI already attaches it).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
resolveVideoCaptureBeyondViewport gated Chrome's beyond-viewport screenshot
path to hardware-GPU captures, to skip the full-surface software
re-rasterization tax. But without beyond-viewport, the viewport-bound
capture clips the bottom edge of any frame containing a native video
surface (the same #1094 tall-portrait guard the alpha capture paths already
hardcode) — leaving ~87 bottom rows black.
This hit two cohorts: software-GPU macOS/Linux hosts, and — worse — EVERY
distributed chunk render, which hardcodes browserGpuMode "software", so the
whole distributed fleet shipped video renders with a black bottom band.
Reporter confirmed forcing resolveVideoCaptureBeyondViewport=true fixes it.
Correct output wins over the software perf optimization: enable
beyond-viewport for any render with a native video surface, regardless of
GPU mode. Drops the now-vestigial browserGpuMode parameter (and its type)
and updates both call sites.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
## Stack
1. **#2298 — DE router stall watchdog** ← you are here
2. #2300 — video bottom-edge clip
3. #2301 — feedback bug-report guidance
## Problem
The DE parallel router auto-enables the interleaved parallel-**streaming** capture for the ≥24 GB macOS trial cohort. If a worker wedges mid-capture (a hung seek/screenshot at an early frame), the render makes **zero frame progress** yet sits until the per-frame CDP `protocolTimeout` (~5 min) fires before the pinned self-verify fallback can run — a silent multi-minute hang shipped to real users.
Reported: stuck at frame 2/2031 for 6+ min, no fallback, until the user manually set `HF_DE_PARALLEL_ROUTER=false` (71 s clean).
## Fix
Add a no-frame-progress watchdog to the parallel branch of `runCaptureStreamingStage`:
- Ticks off `executeParallelCapture`'s progress callback. If no **new** frame lands within `HF_DE_PARALLEL_STALL_MS` (default **60 s** — well under the 5-min protocol timeout, ≫ the 15–32 ms/frame budget), it fires.
- On trip: aborts the **reorder buffer** (so peer workers parked in `waitForFrame` reject instead of deadlocking the `Promise.all` pool) and aborts the pool via a **separate** `AbortController` linked to the parent abort.
- The parent `abortSignal` stays un-aborted, so the orchestrator reads the failure as a generic `capture_error` (not a cancellation) and re-renders on the pinned screenshot path — the same fallback a verify failure already uses.
## Test
- Watchdog trips on no progress → rethrows a stall error (routes to fallback).
- A genuine parent-abort is **not** relabeled as a stall (stays a cancellation).
The DE parallel router auto-enables the interleaved parallel-streaming
capture for the >=24GB macOS trial cohort. If a worker wedges mid-capture
(a hung seek/screenshot at an early frame) the render made no frame
progress yet sat until the per-frame CDP protocolTimeout (~5 min) fired
before the pinned self-verify fallback could run — a silent multi-minute
hang shipped to real users (report: stuck at frame 2/2031 for 6+ min,
no fallback, until HF_DE_PARALLEL_ROUTER=false).
Add a no-frame-progress watchdog to the parallel branch of
runCaptureStreamingStage. It ticks off executeParallelCapture's progress
callback; if no NEW frame lands within HF_DE_PARALLEL_STALL_MS (default
60s, well under protocolTimeout and >> the 15-32ms/frame budget), it
aborts the reorder buffer (unsticking peer workers parked in waitForFrame
so the pool doesn't deadlock) and aborts the pool via a SEPARATE
controller linked to the parent abort. Because the parent abortSignal
stays un-aborted, the orchestrator reads the failure as a generic
capture_error (not a cancellation) and re-renders on the pinned screenshot
path — the same fallback a verify failure already uses.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Addresses review on #2266: the gradient-text exclusion was too broad — any
background-clip:text skipped the invisible-text check, so a broken/missing
gradient (clip:text with no image and a transparent background, which paints
nothing) went unreported. Now exclude only when a real background fills the
glyphs (background-image != none, or an opaque background-color). Expands the
test suite to the reviewer's full case set: direct transparent fill, inherited
transparent fill over an opaque child color, color:transparent fallback, opaque
baseline, gradient-over-real-background exclusion, broken-gradient still flagged,
and empty-text no-op.
The layout-audit test mock returns computed styles as plain camelCase
properties (no getPropertyValue), so the invisible-text detector threw
'cs.getPropertyValue is not a function' and broke the whole audit in CI. Read
webkitTextFillColor/webkitBackgroundClip/backgroundClip by property to match
the rest of the script (works in a real browser too). Adds tests: flags
transparent -webkit-text-fill-color, ignores opaque color, ignores gradient
text (background-clip:text).
Wild report (5th in cluster, CLI 0.7.53): snapshots omitted all text while
check passed — text painting with a transparent -webkit-text-fill-color (which
overrides `color` for the glyph fill AND inherits, so a parent's transparent
fill silently blanks descendant text that has its own opaque `color`) renders
invisible, but every geometry/occlusion/contrast audit missed it. Contrast in
particular reads `color`, not the fill that actually paints, so white-`color`
+ transparent-fill text scored as high-contrast and passed.
Add an invisible-text detector to the layout audit: flag any text element whose
effective fill (computed -webkit-text-fill-color, which already resolves to
`color` when unset) is transparent. Gradient/clipped text (background-clip:text)
legitimately uses a transparent fill and is excluded. Verified: check now fails
on an inherited-transparent-fill fixture (text_not_painted) while gradient text,
body-inherited color, sub-composition color, and real registry examples stay
clean.
Addresses review on #2264: the localization helper had one broad catch around
both dynamic producer resolution and injector execution, so it couldn't tell a
benign 'producer not in this environment' from a real injector/fetch failure,
and emitted no diagnostic. Split into loadFontInjector() (returns null when the
module is absent — silent fail-open) and localizeWithProducer() (warns ONCE per
distinct message when the injector itself throws, then fails open). Per-family
resolution failures remain the injector's own responsibility (producer's
warnUnresolvedFonts). The localizer seam is injectable; tests now cover success,
producer-unavailable, injector-throw, warn dedup, and call-site integration.
The CLI test job builds with --filter '!@hyperframes/producer', and render.ts
imports producer only as a type — so a static import("@hyperframes/producer")
in the font-localization helper failed Vitest's transform-time module
resolution ("Failed to resolve entry for package"), breaking checkBrowser
tests and the helper's own test. Keep the specifier out of the static module
graph (@vite-ignore + variable specifier) so it resolves at runtime only:
production/installed CLI has producer in node_modules and localizes fonts;
the test env fail-opens to the plain bundle. Localizer is now injectable so
the helper's unit tests cover it without needing producer resolvable.
Addresses review feedback on #2254: the previous fix hardcoded every clone
to the full composition width/height regardless of the source scene's own
sizing, silently overriding any authored explicit width/height. Measuring
each scene's getBoundingClientRect() while still live in the document (its
inset:0 already resolves correctly against the real ancestor chain there)
and reapplying that exact box to the clone fixes the 0x0 collapse without
ever clobbering an author's own layout.
Addresses non-blocking review feedback on #2256: the previous shape used
err.message.includes(...) inside a catch to distinguish a real headroom
violation from a statfsSync failure — a future message tweak would silently
fail open (statfs-unsupported and headroom-violation would take the same
code path). Now statfsSync failure returns early (skip the gate) and the
headroom check/throw happens in plain control flow outside any try/catch.
## What
Two fixes from adversarial testing of the DE parallel router (10 hostile comps, routed vs screenshot-baseline PSNR). The router itself held — both bugs are in general drawElement fast capture, and one slipped past self-verify.
### 1. Compile gate: ancestor background-image (`producer`)
`drawElementService`'s per-frame ancestor fill replicates what lies behind the captured subtree by walking up the DOM for the nearest non-transparent **`backgroundColor`**. A background-**image** (`linear-gradient`, `url()`) on `body`/`html`/a wrapper reads as transparent in that scan, so a deeper ancestor's solid color paints instead wherever the subtree leaves pixels uncovered.
Measured repro: body `linear-gradient` + html solid color + an element shrinking late in the comp → DE paints the html purple instead of the body gradient. 30.9 dB min frame vs baseline, visually unmistakable. Identical damage single-worker and parallel — general DE bug, in every wild DE render matching this (very common) authoring pattern.
Fix: `detectAncestorBackgroundImage()` in the compiler (DOM-aware — inline styles on the root's ancestor chain + `<style>` rules resolved via `querySelectorAll`, so class-selected wrappers are covered; backgrounds *inside* the root are deliberately not matched). New compile gate `ancestor_background_image`, same shape as the 3D/mix-blend gates, bypass `HF_FAST_CAPTURE_ANCESTOR_BG=true`.
### 2. Self-verify tail sample (`engine`)
The verify grid sampled at `(i+1)/(k+1)` → [20/40/60/80]% of the timeline. The damage above starts at ~79% and peaks after the last sample — verification **passed** on output that bottomed at 30.9 dB (threshold 32 dB would have caught it, it just never looked there).
Fix: `computeDeVerifySampleFractions()` — first k−1 samples evenly spaced, last pinned at 95%. Default k=4 grid becomes [25/50/75/95]%. Kills the whole late-onset damage class, not just this repro.
## Validation
- Repro comp (body gradient + shrink reveal): now gates → baseline route, 54.7 dB avg vs ground truth (was 41.6 avg / 30.9 min with the purple surround)
- Control comp (nested stacked fades, routed): still routes, verify grid `[90, 180, 270, 342] of 360`, passes, 60.6 dB avg — unchanged
- Full adversarial matrix context: 6/10 comps routed clean (49–68 dB min), blend/3D gated correctly, animated-canvas damage caught by verify at 16.5 dB with clean revert, video comps route legitimately (frames pre-extracted)
- Tests: 7 new detection cases (`htmlCompiler.test.ts`), 5 new grid cases (`frameCapture-verifySampleFractions.test.ts`); `compileStage.test.ts` + `frameCapture.test.ts` suites green
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Two non-blocking review notes from Rames, both addressed:
1. Trial polarity inverted to OPT-IN: disableDeParallelRouterTrial →
enableDeParallelRouterTrial. renderLocal is exported, so any programmatic
consumer (future studio-server path, test harness, distributed runner)
previously inherited the trial and its process-wide env-var/module-latch
state without knowing to disable it — and concurrent invocation races
that state. Now only the CLI's own sequential call sites opt in (the
single top-level render, and batch at concurrency 1); everyone else gets
no trial by default. The doc comment names the sequential-invocation
assumption explicitly.
2. deSelfVerifyFallback semantic narrowing documented at both declarations
(RenderCaptureObservability + RenderPerfSummary.drawElement): since the
pinned-fallback retry was widened, the flag means verify-triggered
SPECIFICALLY — OOM/capture_error fallbacks report false with
deFallbackReason carrying the reason. Dashboards keyed on
de_self_verify_fallback=true as "any fallback fired" must migrate to
de_fallback_reason IS NOT NULL (also called out in the PR body for the
observability rebuild to pick up).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Five findings from a fifth (final scoped) max-effort review of the previous
commit, all local:
1. writeConfig now writes atomically (pid-suffixed temp file + renameSync —
rename within one directory is atomic on POSIX). This closes the real
hazard behind the review's torn-read finding: readConfig's corrupted-file
catch RESETS the config to defaults (telemetry re-enabled, anonymousId
rotated, trial fields wiped), so a concurrent reader catching a
non-atomic write mid-flight would silently destroy the user's config —
and the previous commit's per-render readConfigFresh() at the arm site
multiplied exposure to exactly that window. Verified against a real
filesystem, not just the mocked unit tests.
2. writeConfig now returns whether the write landed (errors still swallowed
— telemetry must never break the CLI). persistDeParallelRouterTrialFired
uses it to stop immediately on a genuine fs failure (retrying an
unwritable file is pointless) and reserve its retries for actual
concurrent clobbers, instead of 3 blind write attempts + 4 disk reads.
3. The persistence-failure console.warn is now !quiet-gated like every
other trial message — a quiet/batch-json render on an unwritable
~/.hyperframes no longer emits unexpected stderr that CI wrappers
asserting empty stderr would misread as a render failure. The in-process
latch already guarantees the safety behavior whether or not the warning
prints.
4. The arm site short-circuits on the in-process fired latch BEFORE the
fresh config read — post-fired batch rows no longer pay a per-row config
read + parse + shared-cache invalidation for an answer module state
already knows.
5. Replaced the new `as T` assertions in render.test.ts's config-state
factory with an explicitly typed vi.hoisted return (repo TypeScript
convention: no `as T`).
config.test.ts: node:fs mock gains renameSync (faithful to the new atomic
write); new test covers the success/failure return and asserts no temp file
survives a write.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three root causes from a fourth max-effort review (15 raw findings deduped;
the synthesize step died on a session limit so they arrived unmerged):
1. The previous commit's telemetryEnabled fix was ineffective: the arm site
passed readConfig() — the process-lifetime cache — into
isDeParallelRouterTrialBlocked, making it exactly as stale as the
shouldTrack() memoization it claimed to bypass. A mid-batch
`hyperframes telemetry off` (or another process persisting fired=true)
was never observed. Now reads readConfigFresh() at the arm site; the
test mock previously hid this because readConfig/readConfigFresh were
behaviorally identical views over one shared object.
2. The verify-and-retry write loop double-counted a render whenever OUR
write landed but a concurrent writer advanced the file before our
verify read — the retry re-applied the increment on top (two renders
→ three counts), tripping the 25-render exposure cap early and
permanently killing the trial with less telemetry than the cap was
designed to allow. Reworked: the render COUNTER is written exactly
once, unverified (a lost increment under-counts by one — benign); only
the FIRED flag is verified and re-asserted, which is idempotent, so
retries can no longer corrupt anything
(persistDeParallelRouterTrialFired).
3. writeConfig swallows all fs errors, so on an unwritable ~/.hyperframes
a reverted outcome could never persist — the trial would re-arm and
re-fail on every subsequent render forever, silently. Added an
in-process fired latch (set at decision time, before persistence is
attempted) consulted by the blocked-check, plus a one-time console
warning when persistence exhausts its attempts. Later processes still
re-arm (disk is the only cross-process channel), but each process now
stops after at most one failure it couldn't record.
Test infrastructure fix enabling all of the above to be tested: the config
mock now models disk vs cache SEPARATELY (readConfig serves the cache,
readConfigFresh re-reads "disk", writeConfig updates both) with a
failWrites hook simulating the real writeConfig's silent error swallowing.
The old single-shared-object mock made cached-vs-fresh mis-routing and
retry iterations untestable by construction.
3 new regression tests: mid-batch opt-out observed through the cache;
fired flag re-asserted after a lost write WITHOUT re-counting the render;
unwritable-config latch blocking re-arm. 56 tests total across
render.test.ts + config.test.ts.
Not fixed (by design): the widened pinned-fallback retry paying a doubled
render on deterministic mid-stream failures (e.g. ENOSPC) — the accepted
tradeoff of the fallback design; cancellation and OOM are special-cased.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Six findings from a third max-effort code review, focused on the previous
commit's fixes:
1. --batch-concurrency N>=2 runs genuinely concurrent renderLocal() calls
(Promise.all workers in batchRender.ts), which can't safely share the
trial's one process-wide env var + module flag — a row finishing first
could tear down the env var/flag mid-render for a sibling row still in
flight. Rather than attempt to make shared process-global state safe
under real concurrency, added RenderOptions.disableDeParallelRouterTrial
and set it whenever batchConcurrency > 1 — the trial simply isn't
offered when it can't be evaluated safely.
2. maybeConsumeDeParallelRouterTrial's "outcome === undefined" no-op guard
almost never fired: aggregateDrawElement (perfSummary.ts) defaults
parallelRouter to the string "none" for every render, whether or not
drawElement/the router ever engaged — never undefined. Every ordinary
render below the router's own frame threshold (the common case) was
ticking the render-count backstop, tripping
DE_PARALLEL_ROUTER_TRIAL_MAX_RENDERS after 25 completely unrelated
renders that never touched the router. Now treats "none" the same as
undefined.
3. isDeParallelRouterTrialBlocked relied solely on shouldTrack(), which
memoizes its verdict once per process — during a long --batch run, a
`hyperframes telemetry off` issued from another terminal mid-batch would
never be observed. Restored a direct config.telemetryEnabled check
(read fresh every call, unlike shouldTrack()'s cache) alongside it.
4. maybeConsumeDeParallelRouterTrial's config write had no way to detect a
losing race against a concurrent process — added a verify-and-retry
loop (write, re-read fresh, retry up to 3x if a concurrent writer
landed in between) that narrows the window further without a full
file-locking rewrite.
5. The trial could arm before the first-run telemetry disclosure
(showTelemetryNotice) was guaranteed to have printed — that notice runs
via a fire-and-forget, unawaited dynamic import in cli.ts with no
ordering guarantee relative to the render command. Rather than touch
that pre-existing async bootstrap chain, gated the trial on
config.telemetryNoticeShown: it simply never offers itself on a fresh
install's very first invocation.
6. Added a dedicated config.test.ts exercising readConfig/readConfigFresh/
writeConfig through the REAL module (node:fs mocked with an in-memory
fake, not a HOME-env hack) — readConfigFresh's cache-bypass and the
type-guarded boolean/number parsing had zero coverage through the real
implementation before this.
Also fixed the test fixture that was supposed to cover finding #2 but used
an unrealistic `drawElement: {}` shape instead of the real
`{ parallelRouter: "none" }` aggregateDrawElement actually produces.
Extracted applyDeParallelRouterOutcome to keep maybeConsumeDeParallelRouterTrial
under the repo's complexity gate after adding the retry loop.
11 new/updated tests in render.test.ts (56 total) + 7 new tests in
config.test.ts. Verified against fallow's audit gate clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Four confirmed findings from a max-effort code review of the CLI trial
mechanism:
1. maybeEnableDeParallelRouterTrial's `process.env.HF_DE_PARALLEL_ROUTER
!== undefined` guard couldn't distinguish "the user set this" from "an
earlier renderLocal() call in this same process already armed it" — so
in --batch (all rows share one process), only row 1's outcome could
ever reach maybeConsumeDeParallelRouterTrial. A revert on any later row
was silently never persisted. Added a module-level
deParallelRouterTrialManagedByUs flag to disambiguate, with a test-only
reset export since it's process-lifetime state a real CLI invocation
never needs to reset but a test suite sharing one module instance does.
2. writeConfig is a non-atomic whole-file overwrite with no locking, and
readConfig's cache never invalidates — a concurrently running second
CLI process (another terminal, a parallel script; doesn't even need to
be a render, any command calls incrementCommandCount) could silently
clobber a just-persisted deParallelRouterTrialFired:true with its own
stale snapshot. Added readConfigFresh (bypasses the cache) and use it
immediately before the trial's read-modify-write, narrowing the race
window without a full config-subsystem locking rewrite.
3. The prior commit's semantics flip removed the only exposure cap — a
healthy router that never reverts now force-enabled the experimental
path on every eligible render forever. Added
DE_PARALLEL_ROUTER_TRIAL_MAX_RENDERS (25) as a backstop: the trial turns
off after this many engaged renders even absent an actual failure.
4. maybeEnableDeParallelRouterTrial only checked config.telemetryEnabled,
not shouldTrack() — so a dev-mode run or a DO_NOT_TRACK/
HYPERFRAMES_NO_TELEMETRY user got the experimental path silently armed
while telemetry was simultaneously blocked underneath it. Now gates on
shouldTrack() (a strict superset).
Also fixed, lower severity: readConfig's deParallelRouterTrialFired/
deParallelRouterTrialRenderCount parsing now validates the JSON type
explicitly instead of a bare truthy/nullish read, so a hand-edited or
corrupted config can't have the string "false" misread as truthy.
Refactored maybeEnableDeParallelRouterTrial into three smaller functions
(isDeParallelRouterTrialBlocked, stopManagingDeParallelRouterTrial) to
bring cyclomatic/cognitive complexity back under the repo's threshold —
also de-duplicates the "stop managing the env var" logic shared with
maybeConsumeDeParallelRouterTrial.
14 new/updated tests (43 total in render.test.ts), including a direct
regression test for the batch re-entrancy scenario and a loop test for the
render-count cap. Verified the config primitives end-to-end against a real
file, not just the mocked unit tests.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- SSOT: the three assemble-index.mjs BGM fallbacks now import
bgmDefaultVolume() from media-use's bgm.mjs instead of duplicating the
0.12/0.9 literals (both reviewers). The cross-skill relative import
matches the existing dependency (each workflow's audio.mjs adapter
already resolves ../../media-use/audio/scripts/audio.mjs).
- STATUS_ROLE_KEY: extended with info|neutral|alert|caution|critical —
same hue-carries-meaning class as the original set (all 3 copies).
- bgm.md: phrase the default as bgmDefaultVolume()/BGM_BED_VOLUME with
"currently 0.12" so the prose survives future tuning.
- fetch-pr.mjs: drop dead mergeCommit field from the gh pr view FIELDS
list (version resolution uses mergedAt only).
- music-to-video assemble-index.mjs: comment documenting why its BGM
stays at 0.8 under VO — music is the content there, not a narration
bed, so the explainer pipelines' 0.12 default deliberately does not
apply.
Not changed: pickAccent's chroma fallback — both call sites pass keyless
capture palettes (tokens.json hex lists), so no status-role keys exist
to filter on; the keyed preset path goes through semanticColors and the
build-frame remix, which this PR already fixed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Behavior fixes surfaced by the prompt-guide validation campaign (Tier 1+2
of the upstream bug list; Tier 3 tracked in #2107). Split out from the
doc-only updates, which follow in a separate PR.
- BGM level: default bed volume under narration was 0.8 linear (~-2 dB,
~16 dB too hot vs voice). Now 0.12 (~-18 dB) via shared
bgmDefaultVolume() in media-use bgm.mjs + assemble-index fallbacks in
faceless-explainer / pr-to-video / product-launch-video. Explicit
volume still wins; silent-film 0.9 and music-to-video unchanged.
Adds bgm.test.mjs (3 cases); bgm.md reference updated to match.
- Caption accent: semanticColors() ranked accents purely by chroma, so a
preserved status red (#dc2626) outranked the brand accent and captions
highlighted in error-red. Status-keyed colors now excluded via shared
STATUS_ROLE_KEY regex consumed by both tokens.mjs and build-frame.mjs
(all three skill copies kept in sync).
- Voice threading: workflow SKILL.md Step 3.1 blocks now instruct
choosing the narration voice from the user's ask and passing
--voice <id>; previously "a male voice" was silently ignored and the
default (Marcia/am_michael) always won.
- fetch-pr shipping version: MERGED PRs get best-effort shipped_version
+ version_source in pr.json (first release published at/after merge,
else default-branch package.json marked unreleased); ingest surfaces
it as a 'Shipped in:' brief line; story-design.md forbids inventing
versions when absent.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Only consuming telemetry from one data point per install badly undersampled
the "routed" (successful) outcome — the far more common case. Changed
maybeConsumeDeParallelRouterTrial to only turn the trial off when the
router's OWN safety net actually fired (deParallelRouter === "reverted"),
not on a clean "routed" success. This runs the experiment on every eligible
render for an install indefinitely until it hits one real failure, then
stops for that install going forward — trading a slightly higher per-install
ceiling on experimental-path exposure for dramatically more successful-
routing telemetry volume across the fleet.
Also fixed a related edge case while updating this: a render that merely
"routed" (router fired, self-verify never even tripped) but then crashed
for an unrelated reason (e.g. cancellation) no longer counts as a router
failure — only "reverted" (the router's fallback path actually engaged)
does. Cancelling a render isn't evidence the router is unsafe.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
HF_DE_PARALLEL_ROUTER is a producer env var with no self-serve opt-in path
for real users, so waiting for someone to manually enable it would never
produce the real-traffic telemetry (revert rate, verify-db distribution)
the router's soak plan calls for.
renderLocal now enables the experiment for free on a fresh install's CLI
renders until it actually engages once (routed or reverted — either
produces telemetry), then persists that to ~/.hyperframes/config.json and
never touches it again for that install. A render whose frame count never
crosses the router's own eligibility threshold doesn't consume the trial —
it stays available for a later render that does qualify.
Never overrides a user's own explicit HF_DE_PARALLEL_ROUTER setting, and
only engages when telemetry is enabled (no point risking the experimental
path if we can't record the resulting signal). Scoped to the in-process CLI
render path only — Docker renders don't thread perfSummary/errorDetails
back to the CLI process, so trial consumption can't be detected there.
Verified the config round-trip against a real file (fresh install ->
undefined -> write true -> persists across reread), not just the mocked
unit tests.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Three defects found by max-effort code review of this branch:
1. The Bun OOM exact-match regex was defeated by this codebase's own
parallel-worker error wrapping. executeParallelCapture/formatWorkerFailure
(parallelCoordinator.ts) always wrap a worker's error as
"Worker N: <message>", optionally suffixed and joined with other workers'
segments, all prefixed "[Parallel] Capture failed: ". That wrapping
defeated the exact-message check for exactly the cohort (deParallelRouter
routed, N separate Chrome processes) the OOM-drops-to-1 fix targets — a
real OOM there would retry at the SAME worker count instead of dropping
to 1. Added a second pattern that recovers the signal by requiring
"out of memory" appear as the WHOLE content of a "Worker N: ..." segment
(bounded by end-of-string/"; "), preserving the same exact-match property
(no bare substring match) while surviving the wrapping. Verified against
the real wrapping logic, not a hand-typed guess at its shape.
2. shouldRetryViaPinnedFallback didn't exclude cancellation, so aborting a
render mid-capture on the pinned router/inversion cohort would detour
through spawning a fresh encoder/capture session before the outer catch's
RenderCancelledError branch ended the render — delaying "stop" with a
pointless resource spin-up/tear-down. Added an isCancellation param
(checked first, before isVerifyError) using the same
`err instanceof RenderCancelledError || abortSignal?.aborted` check the
outer catch already uses.
3. deFallbackReason (this PR's new "oom"/"capture_error" values) was set
locally but never mirrored into RenderCaptureObservability alongside
deSelfVerifyFallback, so a render that fails AFTER a fallback attempt
(perfSummary never built) was indistinguishable in render_error telemetry
from one that never attempted any fallback — undercutting the "how often
does the OOM retry fire on a render that still ultimately fails"
question this branch exists to answer. Threaded through
RenderCaptureObservability → RenderObservabilityTelemetryPayload →
renderObservabilityTelemetryPayload, mirroring the existing
deSelfVerifyFallback plumbing.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The old setup section was one long token-minting flow with the MCP
connector as an afterthought — but motion/shaders/storyboards need
only the connector (no token at all), and even brand tokens are
easier via MCP on any non-Enterprise plan. A reader wanting only
motion had to wade past REST scope tables meant for a different path.
Restructured into a decision table (what you want → which credential)
followed by two equal, independent steps. Step A's scope list now
states the 3 boxes to check on a normal (non-Enterprise) account
up front, instead of a generic 4-scope table the reader has to
interpret themselves. Step B now states the MCP connector's actual
capability (variable reads work on any plan, rate-limited by tier —
6 calls/month on Starter) instead of "no scopes to configure," which
undersold what it can do.
Co-Authored-By: Claude Opus <noreply@anthropic.com>
Miga's re-review caught it: the new RATE_LIMITED row (client auto-
retries with backoff) landed alongside the old pre-retry row ("wait a
minute and retry; chunk batch renders"), leaving two rows for the same
code — one accurate, one stale. Keep only the current one.
Co-Authored-By: Claude Opus <noreply@anthropic.com>
Two real incidents this week had agents skip /figma entirely and drive
Figma via raw MCP tools (get_metadata/get_screenshot/get_design_context)
when a figma.com URL landed inside a creation-workflow skill. Root cause:
none of the creation workflows mention Figma at all, and the only routing
table that does (/hyperframes) is skipped whenever a workflow is invoked
directly rather than through the entry router — which is the common path.
Going raw loses real infrastructure the CLI/skill guarantees: sanitizeSvg()
before freezing (raw-fetched SVGs are unsanitized), .media/manifest.jsonl
provenance (no cache-hit, no version tracking), and brand-token var()
binding (colors bake as literals, so a later Figma brand change can't
propagate without a full re-import).
Added a "figma source" callout to every creation workflow that could
plausibly receive a figma.com link (product-launch-video, website-to-video,
general-video, motion-graphics, slideshow), plus a defense-in-depth line
in /hyperframes's own routing checklist. The fix lives in the workflows
themselves so it doesn't depend on the entry router being consulted.
Co-Authored-By: Claude Opus <noreply@anthropic.com>
Found while testing the previous commit's OOM-drops-to-1-worker fallback
end-to-end: the producer's deployed runtime is Bun (JavaScriptCore), not
Node (V8) — see packages/gcp-cloud-run/Dockerfile's `bun dist/server.js`
entrypoint. All 7 MEMORY_EXHAUSTION_ERROR_PATTERNS are V8-specific allocation
failure signatures; JSC's equivalent for the same single-oversized-allocation
RangeErrors is the bare string "Out of memory" (verified against real Bun
behavior), which none of them match. Without this, isMemoryExhaustionError
returns false for genuine production OOM, so the memory-specific worker-count
reduction just added would never actually engage where it's deployed — every
OOM would fall through to the generic capture_error retry path instead.
Matches the FULL (trimmed) message only, not merely a substring — same
rationale as the existing V8 patterns' comment: "out of memory" also appears
in benign WebGL/GPU console noise that must not trip this classifier.
Verified end-to-end from a script inside the producer workspace (importing
the real @hyperframes/engine source, not a stale globally-cached npm dist a
script outside the workspace would otherwise resolve to): a genuine Bun
RangeError from new Uint8Array(Number.MAX_SAFE_INTEGER) now correctly
classifies as memory exhaustion and drives both resolveInversionRetryPlan
and resolveParallelRouterRetryPlan down to workerCount=1 on retry.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
nodeToHtml routed rasterize eligibility off node.type alone, so a
RECTANGLE/FRAME with an IMAGE fill fell through to the generic <div>
path — fillCss() has no IMAGE case, so it rendered an empty box.
IMAGE-filled nodes now route to rasterize like vectors, regardless of
node.type.
Rasterized nodes (vectors, now image fills too) were also getting
their own fill/corner-radius CSS applied on top of the already-
rendered <img> — a flat color block behind/around the real art,
flattening non-rectangular shapes into rounded rects. decorationCss
now skips background and corner-radius/clip for rasterized nodes;
opacity and effects still apply since those aren't baked into the
export.
tokens.ts's styles-fallback path hardcoded entries: [] regardless of
how many published styles were actually found, so the CLI printed
"recorded published style metadata instead" even when styles()
returned zero results. Added styleCount to the result so the message
reflects what happened, and points at the MCP get_variable_defs
fallback when there's nothing to fall back to.
Co-Authored-By: Claude Opus <noreply@anthropic.com>
shouldRetryViaPinnedFallback retrying OOM was only half the fix: it reused
preInversionWorkerCount/preRouterWorkerCount unmodified, which is
calibration's own pick and can be >= the pinned count that just OOM'd
(calibration wanting 5 while the router pinned to 3). Retrying at equal or
higher parallelism than the failure isn't a remedy — it's the same bet
again, and worsens the odds for this render and anything sharing the host
(PRODUCER_MAX_CONCURRENT_RENDERS runs concurrent jobs in one process).
resolveInversionRetryPlan/resolveParallelRouterRetryPlan now drop to
workerCount=1 specifically when the retry is OOM-triggered — one Chrome
page, the leanest configuration available, not just a different capture
mode at the same worker count. Ordinary self-verify (blank/PSNR) retries
are unaffected — those aren't memory-related, so they keep the
pre-inversion/pre-router count as before.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The router/inversion pin a fixed worker count regardless of calibration —
exactly the scenario a host-contention timeout or worker crash is most
likely under. Previously only a drawElement self-verify failure (blank
frame / PSNR breach) triggered the existing fallback to the calibrated,
non-DE parallel-screenshot path; any other capture-stage failure on a
pinned render just hard-failed the whole job instead of reusing that same
tested safety net.
shouldRetryViaPinnedFallback widens the retry to any capture failure while
deWorkerInversion="inverted" or deParallelRouter="routed", excluding OOM
(the fallback's worker count can be >= the pinned count, so retrying would
likely just OOM again — fail fast instead).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
CodeQL: shell command built from environment values — the oxfmt
invocation interpolated a filesystem-derived absolute path into a shell
string. execFileSync with array args avoids the shell entirely.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Addresses @miguel-heygen's review on #2112:
- BLOCKER: docs/guides/figma.mdx now matches the shipped code/skill —
adds the Library content: Read-only scope row (+ corrects the
'falls back, expected' line that was false without it), and the
troubleshooting table now says bad PATs surface as 403 Invalid token
(not 401), names the scope in FORBIDDEN, and documents RATE_LIMITED retry.
- nit: the batch summary line no longer claims '1 figma request' when every
node was a cache hit — says 'all reused from cache — no figma request'.
- nit: index.md regen moved to a finally, so a mid-batch RENDER_FAILED
leaves index.md consistent with the nodes that did freeze.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
render_error previously carried zero DE-cohort context — a hard failure while
routed (worker crash, OOM, capture timeout from the fixed 3-worker pin
overriding calibration) was indistinguishable from any other failure. The
data existed (RenderCaptureObservability is mutated live and survives into
job.errorDetails on the failure path) but was never projected into the
render_error payload, which only ever drew de_* fields from perfSummary
(success-only).
- RenderCaptureObservability now also records dePreInversionWorkers /
dePreRouterWorkers — the worker count calibration would have picked absent
the experiment — so a resource-pressure failure can be correlated with the
router overriding a lower calibrated count.
- New capture-sourced de_* fields on RenderObservabilityTelemetryPayload,
shared by trackRenderComplete and trackRenderError. Explicit
perfSummary-sourced fields still win on render_complete (spread moved
first in the event object) — this is purely a failure-path fallback.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Rames's inline findings on #2112:
- forbiddenError now RETURNS in every branch (BAD_TOKEN no longer throws
inside) so the caller's single throw covers all cases — no mixed
throw/return contract for a future wrapping caller.
- retryAfterMs capped at 60s: a spec-legal Retry-After: 3600 no longer
silently blocks the CLI for an hour before RATE_LIMITED.
- asset ref gathering extracted to gatherAssetRefs() and made URL-safe:
bare fileKey:nodeId tokens comma-split, but a figma URL with commas in
its query (multi-select node-id=1:2,3:4) is kept whole.
- Documented in SKILL that 429 retry lives in the shared request path, so
EVERY read endpoint retries (not just asset) — blast-radius note the
reviewer asked for. variables intentionally still retries: its fallback
is REQUIRES_ENTERPRISE-only, and a 429 there is transient, not a gate.
Tests: retry-cap (3600→60000), non-styles endpoint retry, gatherAssetRefs
URL-vs-bare split. client 24, cli asset 11.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CI typecheck caught the tokens.test mock missing the new renderNodes
member on FigmaClient (asset/component mocks were updated, this one was
missed).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Extends the scope+retry work from the figma bug-bash (valid report:
9-bugs-with-repros; the skill-not-used report was discarded).
- 403-body parse (bug 4): figma returns 403 {"err":"Invalid token"} for bad
PATs (NOT 401), and 403 {"err":"Invalid scope(s)… requires X"} for missing
scopes. get() now reads the body: "Invalid token" reclassifies to BAD_TOKEN
with re-mint advice; a scope body surfaces figma's own diagnosis verbatim;
else falls back to the endpoint's scope hint. Reads both err and message
(variables endpoint uses message). One fix, honest messages for bugs 1/4/9.
- Batch asset fetch (requested): figma asset accepts multiple refs
(space-separated or comma-joined) of one file and renders them in a SINGLE
/v1/images call via new client.renderNodes — figma's documented per-minute
rate-limit workaround. runAssetImport delegates to runAssetImportMany;
cache-checks per node, batches only the misses, one index.md regen.
- NO_TOKEN box (bug 8): errorBox indented only the first hint line, mangling
the numbered setup list. Indent every line; single-line hints unchanged.
Verified live: 3 refs -> 3 imports -> 1 request; bad token -> BAD_TOKEN not
scope advice. Client suite 22, cli figma 33.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two bugs from live figma-integration use:
1. `tokens` styles fallback 403s on non-Enterprise. /v1/files/:key/styles
needs library_content:read — a scope the setup docs and the generic
FORBIDDEN message both omitted, so the user saw "missing a read scope"
with no way to know which. Each endpoint now carries a scope hint; the
403 names the exact scope (styles → library_content:read). Setup text and
skill scope list updated to include Library content: Read-only.
2. `asset` (and every per-node component render) had no 429 handling — the
message said "back off and retry" but the client didn't. Two imports in
a row tripped the per-minute limit and hard-failed. get() now retries 429
with exponential backoff, honoring Retry-After when present, before
surfacing RATE_LIMITED after maxRetries (default 3). sleep is injectable
so tests don't wait.
Batch multi-node asset syntax (the documented /v1/images comma-ids rate
workaround) is a separate enhancement — retry makes the reported failure
self-heal, including the many-node component path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>