Addresses review on #2266: the gradient-text exclusion was too broad — any
background-clip:text skipped the invisible-text check, so a broken/missing
gradient (clip:text with no image and a transparent background, which paints
nothing) went unreported. Now exclude only when a real background fills the
glyphs (background-image != none, or an opaque background-color). Expands the
test suite to the reviewer's full case set: direct transparent fill, inherited
transparent fill over an opaque child color, color:transparent fallback, opaque
baseline, gradient-over-real-background exclusion, broken-gradient still flagged,
and empty-text no-op.
The layout-audit test mock returns computed styles as plain camelCase
properties (no getPropertyValue), so the invisible-text detector threw
'cs.getPropertyValue is not a function' and broke the whole audit in CI. Read
webkitTextFillColor/webkitBackgroundClip/backgroundClip by property to match
the rest of the script (works in a real browser too). Adds tests: flags
transparent -webkit-text-fill-color, ignores opaque color, ignores gradient
text (background-clip:text).
Wild report (5th in cluster, CLI 0.7.53): snapshots omitted all text while
check passed — text painting with a transparent -webkit-text-fill-color (which
overrides `color` for the glyph fill AND inherits, so a parent's transparent
fill silently blanks descendant text that has its own opaque `color`) renders
invisible, but every geometry/occlusion/contrast audit missed it. Contrast in
particular reads `color`, not the fill that actually paints, so white-`color`
+ transparent-fill text scored as high-contrast and passed.
Add an invisible-text detector to the layout audit: flag any text element whose
effective fill (computed -webkit-text-fill-color, which already resolves to
`color` when unset) is transparent. Gradient/clipped text (background-clip:text)
legitimately uses a transparent fill and is excluded. Verified: check now fails
on an inherited-transparent-fill fixture (text_not_painted) while gradient text,
body-inherited color, sub-composition color, and real registry examples stay
clean.
Addresses review on #2264: the localization helper had one broad catch around
both dynamic producer resolution and injector execution, so it couldn't tell a
benign 'producer not in this environment' from a real injector/fetch failure,
and emitted no diagnostic. Split into loadFontInjector() (returns null when the
module is absent — silent fail-open) and localizeWithProducer() (warns ONCE per
distinct message when the injector itself throws, then fails open). Per-family
resolution failures remain the injector's own responsibility (producer's
warnUnresolvedFonts). The localizer seam is injectable; tests now cover success,
producer-unavailable, injector-throw, warn dedup, and call-site integration.
The CLI test job builds with --filter '!@hyperframes/producer', and render.ts
imports producer only as a type — so a static import("@hyperframes/producer")
in the font-localization helper failed Vitest's transform-time module
resolution ("Failed to resolve entry for package"), breaking checkBrowser
tests and the helper's own test. Keep the specifier out of the static module
graph (@vite-ignore + variable specifier) so it resolves at runtime only:
production/installed CLI has producer in node_modules and localizes fonts;
the test env fail-opens to the plain bundle. Localizer is now injectable so
the helper's unit tests cover it without needing producer resolvable.
* fix(telemetry): expose stalled render stages
* fix(telemetry): preserve capture data on terminal stage events
* fix(telemetry): fix calibration TDZ crash, tag encode/assemble, extend heartbeat cadence
capture_calibration referenced captureStageObservationData before its
declaration (later in the same scope), which would throw a ReferenceError
for any render hitting the calibration path. Hoist the closure and split
workerCount's declaration from its resolution so calibration can safely
read it as undefined before capture strategy resolves worker count.
Also address the two non-blocking review items: wire encode/assemble
stages through captureStageObservationData for consistent tagging, and
extend the heartbeat schedule to repeat every 120s after the initial
30/60/120s ramp instead of going dark on stalls beyond two minutes.
Two non-blocking review notes from Rames, both addressed:
1. Trial polarity inverted to OPT-IN: disableDeParallelRouterTrial →
enableDeParallelRouterTrial. renderLocal is exported, so any programmatic
consumer (future studio-server path, test harness, distributed runner)
previously inherited the trial and its process-wide env-var/module-latch
state without knowing to disable it — and concurrent invocation races
that state. Now only the CLI's own sequential call sites opt in (the
single top-level render, and batch at concurrency 1); everyone else gets
no trial by default. The doc comment names the sequential-invocation
assumption explicitly.
2. deSelfVerifyFallback semantic narrowing documented at both declarations
(RenderCaptureObservability + RenderPerfSummary.drawElement): since the
pinned-fallback retry was widened, the flag means verify-triggered
SPECIFICALLY — OOM/capture_error fallbacks report false with
deFallbackReason carrying the reason. Dashboards keyed on
de_self_verify_fallback=true as "any fallback fired" must migrate to
de_fallback_reason IS NOT NULL (also called out in the PR body for the
observability rebuild to pick up).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`hyperframes skills update` failed hard or looped forever once a skill was
retired/renamed upstream while still installed locally (hyperframes-media folded
into media-use; hyperframes-captions/compose/tts consolidated earlier). Two
paths dead-ended:
- Install: target selection could trust a stale local skills-manifest.json
(findRepoManifest) while `skills add` always installs from the canonical repo.
isCoreSkill matches the `hyperframes-` prefix, so a retired skill was forced
into the target set, `skills add` silently declined it (exit 0), and strict
verifyInstalled threw "Skill(s) still missing after install".
- Prune: upstream `skills remove` scans on-disk directories, so a lock entry
retired before it ever shipped a bundle has nothing to match — a silent
exit-0 no-op that never clears the lock, so detectRemoved re-flags it on
every run.
The stale-skills nudge compounded it: it fired even from `skills update` itself
(pointing users back at the failing command) and its count ignored the removed
bucket.
Resolve update targets against the canonical manifest (checkSkills({ canonical:
true })) so a retired skill is never targeted. Add pruneOrphanedLockEntries to
clear the orphaned lock entries the upstream remover can't (idempotent, so a
second run is a clean no-op). Exclude `skills` from the update-nudge gate and
thread the removed count through the nudge total.
The pipeline already carried FrameCheckOptions; only the flag was
boolean, which meant a pipeline caller tuning severity or seek points
would have them silently dropped — the two sides only agreed because
today's caller happens to match the defaults. Bare --frame-check keeps
the defaults; the value form mirrors --caption-zone's grammar, freezing
the contract before a release pins it.
Port validate's per-media-element clip audit into check's session
(clip_media_fit findings): an intrinsic duration meaningfully shorter
than the data-duration slot silently shortens the slot at render time,
and neither lint nor the runtime listeners can see it. A linter crash
now reports as check_lint_failure instead of masquerading as a runtime
failure. Finding-crop capture failures stay non-gating but emit a
stderr note and a telemetry error event so rollouts can measure the
second-session failure rate.
Five findings from a fifth (final scoped) max-effort review of the previous
commit, all local:
1. writeConfig now writes atomically (pid-suffixed temp file + renameSync —
rename within one directory is atomic on POSIX). This closes the real
hazard behind the review's torn-read finding: readConfig's corrupted-file
catch RESETS the config to defaults (telemetry re-enabled, anonymousId
rotated, trial fields wiped), so a concurrent reader catching a
non-atomic write mid-flight would silently destroy the user's config —
and the previous commit's per-render readConfigFresh() at the arm site
multiplied exposure to exactly that window. Verified against a real
filesystem, not just the mocked unit tests.
2. writeConfig now returns whether the write landed (errors still swallowed
— telemetry must never break the CLI). persistDeParallelRouterTrialFired
uses it to stop immediately on a genuine fs failure (retrying an
unwritable file is pointless) and reserve its retries for actual
concurrent clobbers, instead of 3 blind write attempts + 4 disk reads.
3. The persistence-failure console.warn is now !quiet-gated like every
other trial message — a quiet/batch-json render on an unwritable
~/.hyperframes no longer emits unexpected stderr that CI wrappers
asserting empty stderr would misread as a render failure. The in-process
latch already guarantees the safety behavior whether or not the warning
prints.
4. The arm site short-circuits on the in-process fired latch BEFORE the
fresh config read — post-fired batch rows no longer pay a per-row config
read + parse + shared-cache invalidation for an answer module state
already knows.
5. Replaced the new `as T` assertions in render.test.ts's config-state
factory with an explicitly typed vi.hoisted return (repo TypeScript
convention: no `as T`).
config.test.ts: node:fs mock gains renameSync (faithful to the new atomic
write); new test covers the success/failure return and asserts no temp file
survives a write.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three root causes from a fourth max-effort review (15 raw findings deduped;
the synthesize step died on a session limit so they arrived unmerged):
1. The previous commit's telemetryEnabled fix was ineffective: the arm site
passed readConfig() — the process-lifetime cache — into
isDeParallelRouterTrialBlocked, making it exactly as stale as the
shouldTrack() memoization it claimed to bypass. A mid-batch
`hyperframes telemetry off` (or another process persisting fired=true)
was never observed. Now reads readConfigFresh() at the arm site; the
test mock previously hid this because readConfig/readConfigFresh were
behaviorally identical views over one shared object.
2. The verify-and-retry write loop double-counted a render whenever OUR
write landed but a concurrent writer advanced the file before our
verify read — the retry re-applied the increment on top (two renders
→ three counts), tripping the 25-render exposure cap early and
permanently killing the trial with less telemetry than the cap was
designed to allow. Reworked: the render COUNTER is written exactly
once, unverified (a lost increment under-counts by one — benign); only
the FIRED flag is verified and re-asserted, which is idempotent, so
retries can no longer corrupt anything
(persistDeParallelRouterTrialFired).
3. writeConfig swallows all fs errors, so on an unwritable ~/.hyperframes
a reverted outcome could never persist — the trial would re-arm and
re-fail on every subsequent render forever, silently. Added an
in-process fired latch (set at decision time, before persistence is
attempted) consulted by the blocked-check, plus a one-time console
warning when persistence exhausts its attempts. Later processes still
re-arm (disk is the only cross-process channel), but each process now
stops after at most one failure it couldn't record.
Test infrastructure fix enabling all of the above to be tested: the config
mock now models disk vs cache SEPARATELY (readConfig serves the cache,
readConfigFresh re-reads "disk", writeConfig updates both) with a
failWrites hook simulating the real writeConfig's silent error swallowing.
The old single-shared-object mock made cached-vs-fresh mis-routing and
retry iterations untestable by construction.
3 new regression tests: mid-batch opt-out observed through the cache;
fired flag re-asserted after a lost write WITHOUT re-counting the render;
unwritable-config latch blocking re-arm. 56 tests total across
render.test.ts + config.test.ts.
Not fixed (by design): the widened pinned-fallback retry paying a doubled
render on deterministic mid-stream failures (e.g. ENOSPC) — the accepted
tradeoff of the fallback design; cancellation and OOM are special-cased.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(core): escape NUL delimiters in HFMASK mask token and restore regex
Raw 0x00 bytes in the maskInertRegions token and restore regex made
timingCompiler.ts binary to git and shipped raw NULs into dist/cli.js.
Bun's transpiler (<= 1.3.11) corrupts raw NULs in regex literals into
literal backslash-uFFFD text, so restore never matched: every masked
<style>/<script> region was dropped, the player never initialized, and
bunx renders produced blank white frames showing HFMASK tokens.
Use \u0000 escapes instead, which survive any transpile layer, and add
a byte-level regression test (behavior is identical under Node, so only
a byte check catches this).
Fixes the first half of #2139.
* fix(cli): use NTFS junctions for studio project links on Windows
linkProjectIntoStudioData called symlinkSync(dir, path, "dir"), which
needs Developer Mode or elevation on Windows, so preview and dev in
local-studio mode died with EPERM for default-configured users.
Junctions need no privilege, work for directories, and keep the live
write-back the studio depends on (a copy fallback would decouple the
studio from the real project). Covers both preview and dev, which share
the helper.
Fixes the second half of #2139.
Running the CLI from source (tsx dev script, as CI's smoke job does)
transpiles with keepNames, which rewrites named inner functions in
serialized page closures into __name(...) calls — a helper that exists
in the Node bundle but not in the browser realm. The unified seek
closure was the first validate-path page function with named inner
functions, so 'hyperframes validate' threw '__name is not defined' in
CI while the dist build worked. Every session opener now installs a
no-op __name shim via evaluateOnNewDocument before any page script
runs, immunizing all serialized closures regardless of build mode.
Three filters keep the escalated contrast gate honest, each surfaced by
running check across the registry examples:
- data-layout-ignore (the layout audit's existing decorative opt-out)
now also excludes set-dressing text from the contrast audit — one
vocabulary for 'not copy a viewer must read'.
- Text that has (nearly) left the canvas is skipped: sampling a clamped
off-canvas box reads border pixels and produced the classic false
white-on-white (a cursor exiting the frame).
- Contrast failures follow the same persistence rule as layout findings:
observed at a single sample of a multi-sample sweep demotes to
warning; held failures gate.
Example fixes the sweep exposed: motion-blur's 21 deliberately-dim rail
labels are marked decorative (its vivid labels pass on their own);
nyt-graph's subtitle and source-note adopt the gate's suggested
compliant gray; decision-tree's declared duration drops 15s to 10s —
its content ends at ~9.5s and every render shipped a blank white tail.
Every duplicated decision gets one owner: rectToBbox lives in checkTypes
(was verbatim in pipeline and browser layers); the audit seek tuning is
one exported AUDIT_SEEK_OPTIONS consumed by check and the deprecated
inspect path; zoom padding/scale defaults export from the capture module
instead of re-literalized in three files; the optional run_id property
is built by one helper across all three telemetry events; check's
--max-transition-samples parsing reuses its own positiveInteger helper;
validate drops a leftover re-export and redundant explicit-default args
go away.
Tests: the contrast candidate round-trip gains a real integration
anchor (the actual browser script eval'd in-page, a wrapper asserting
finish receives the page-script bbox shape) replacing regex-over-source
as the primary guard; the redundant geometry source-golden and a
duplicated deprecation-envelope assertion are dropped.
Scaffolded projects' npm run check now invokes the single check command
instead of chaining lint, validate, and inspect (three Chrome boots
become one). The CLI skill, its correctness reference, the entry skill's
capability map, README/docs catalog rows, the Mintlify CLI page (new
check section, deprecation banner on inspect), template CLAUDE/AGENTS
(byte-identical), root CLAUDE/AGENTS, and every creation-workflow skill
that taught the old sequence all point at check. snapshot keeps its
standalone sections; validate/inspect stay documented as deprecated
aliases with their check equivalents.
Layout findings now distinguish held defects from entrance/exit
transients: a dynamic issue seen at a single grid sample demotes to
info, while content_overlap held across two-plus samples (or 500ms+)
promotes to error, resolving the long-standing re-promotion TODO. Static
compositions keep their severity. check gains a sweep_static error when
a 3s+ composition shows zero geometry change across every sample (a
frozen timeline makes every green verdict unreliable); skipped when the
motion sidecar already reported motion_frozen. text_occluded findings
carry a coveredFraction; atomic labels (short, no whitespace) flag on
any cover while prose needs 15%, since partial cover changes what a
short label reads as.
Deprecation-test scaffolding consolidates into deprecationTestHarness;
tier logic and logger tests restructured under the complexity gate
without suppression markers.
Detection mechanics adapted from Adam Rosler's open-sourced
visual-linter design (github.com/Adam-Rosler/hyperframes-visual-linter-design);
the elementFromPoint paint model, opt-out attributes, and single-audit
architecture are unchanged.
snapshot --zoom <selector|x,y,w,h> + --zoom-scale (default 3) crops via
Puppeteer clip at raised deviceScaleFactor — density changes, layout
never does. Selector resolves per frame with 24px padding; no match is
a loud error, and a frame whose clamped region is a sliver (element
collapsed or animated off-canvas) is skipped with a stderr note rather
than written as a useless few-pixel image.
check --snapshots additionally writes finding-NN-<code>.png crops for
error findings with bboxes (cap 12, deterministic re-seek in a second
session) and draws labeled annotation boxes on overview frames via a
transient overlay injected only after audits complete. Skill reference
gains the zoom workflow: check reports a finding, zoom into it, fix,
re-check.
One stderr notice per invocation and _meta.deprecated: true in JSON mode
(shared helper next to withMeta; layout owns both inspect and layout via
createInspectCommand). Help descriptions gain the pointer. No behavior
change; removal ships separately once migration telemetry says usage
has decayed.
fix(producer): route info/debug logs to stderr — the compiler's
'Localized remote media' line was landing on stdout ahead of validate's
--json payload, breaking every piped consumer. Diagnostics now share
stderr with warn/error; render progress uses its own channel.
HYPERFRAMES_RUN_ID (trimmed, 128-char cap) attaches as run_id to the
generic cli_command / cli_command_result events, absent when unset, so
an orchestrator setting it per design element can group a verify loop's
invocations in analytics. check additionally emits one check_report
event per invocation (including lint-short-circuited and failing runs):
gate booleans, per-class error/warning counts, launch/seek/contrast
phase timings, sample counts, ok and exit code. Timings stay internal;
no command output changes.
Ports the EF bridge's captionZone and frameCheck semantics as opt-in
flags so the bespoke bridge can be retired: --caption-zone takes
fractional band geometry (x0;y0;x1;y1) with optional severity routing
and seek points, defaults matching the bridge (caption seek [1], frame
seek [0.5], 2px tolerance, 0.05 opacity floor, 4px minimum size, 0.95
full-frame exclusion, center-in-band comparison, tag|text dedup).
--frame-check adds media bounds detection (img/svg/video/canvas) the
always-on text canvas_overflow never covered, reusing overflowFor.
Breach floor: max(120px, 6% of min canvas dimension). Band math derives
from the composition's own canvas, portrait included. Both gates off by
default; plain check output unchanged.
One command, one Chrome boot: in-process lint gate (browser skipped on
lint errors), passive runtime capture wired before navigation, layout +
motion + contrast audits over one seek grid, optional --snapshots
persisting the contrast-pass screenshots. Aggregated --json envelope
{ok, lint, runtime, layout, motion, contrast, snapshots}; findings carry
selector/data-*/source-file/bbox/time anchors, contrast findings include
fg/bg, measured vs required ratio, and a compliant color suggestion.
Contrast AA failures gate the exit code (they were warning-only in
validate); --strict gates warnings.
Contrast candidates round-trip verbatim between __contrastAuditPrepare
and __contrastAuditFinish: the page script owns their shape (bbox w/h),
and normalizing them Node-side made every sample rect NaN — the audit
reported zero checked elements as green. Regression-pinned in
check.test.ts; E2E on a low-contrast fixture now exits 1 with 8 findings.
Measured on kinetic-type: check 5.6s vs 23.0s for sequential
validate + inspect + snapshot.
seekCompositionTimeline becomes the single seek implementation with
per-caller settle options (rAF mode, font wait, settle sleep), replacing
the divergent local seekTo copies in validate and layout. All three
launch paths now build args via the engine's buildChromeArgs; screenshot
paths keep the engine's software-GPU default for deterministic output.
inspect gains one transient content_overlap warning on product-promo
(t=12.22s): the gsap.ticker.tick flush samples timeline state the old
layout seek missed.
Six findings from a third max-effort code review, focused on the previous
commit's fixes:
1. --batch-concurrency N>=2 runs genuinely concurrent renderLocal() calls
(Promise.all workers in batchRender.ts), which can't safely share the
trial's one process-wide env var + module flag — a row finishing first
could tear down the env var/flag mid-render for a sibling row still in
flight. Rather than attempt to make shared process-global state safe
under real concurrency, added RenderOptions.disableDeParallelRouterTrial
and set it whenever batchConcurrency > 1 — the trial simply isn't
offered when it can't be evaluated safely.
2. maybeConsumeDeParallelRouterTrial's "outcome === undefined" no-op guard
almost never fired: aggregateDrawElement (perfSummary.ts) defaults
parallelRouter to the string "none" for every render, whether or not
drawElement/the router ever engaged — never undefined. Every ordinary
render below the router's own frame threshold (the common case) was
ticking the render-count backstop, tripping
DE_PARALLEL_ROUTER_TRIAL_MAX_RENDERS after 25 completely unrelated
renders that never touched the router. Now treats "none" the same as
undefined.
3. isDeParallelRouterTrialBlocked relied solely on shouldTrack(), which
memoizes its verdict once per process — during a long --batch run, a
`hyperframes telemetry off` issued from another terminal mid-batch would
never be observed. Restored a direct config.telemetryEnabled check
(read fresh every call, unlike shouldTrack()'s cache) alongside it.
4. maybeConsumeDeParallelRouterTrial's config write had no way to detect a
losing race against a concurrent process — added a verify-and-retry
loop (write, re-read fresh, retry up to 3x if a concurrent writer
landed in between) that narrows the window further without a full
file-locking rewrite.
5. The trial could arm before the first-run telemetry disclosure
(showTelemetryNotice) was guaranteed to have printed — that notice runs
via a fire-and-forget, unawaited dynamic import in cli.ts with no
ordering guarantee relative to the render command. Rather than touch
that pre-existing async bootstrap chain, gated the trial on
config.telemetryNoticeShown: it simply never offers itself on a fresh
install's very first invocation.
6. Added a dedicated config.test.ts exercising readConfig/readConfigFresh/
writeConfig through the REAL module (node:fs mocked with an in-memory
fake, not a HOME-env hack) — readConfigFresh's cache-bypass and the
type-guarded boolean/number parsing had zero coverage through the real
implementation before this.
Also fixed the test fixture that was supposed to cover finding #2 but used
an unrealistic `drawElement: {}` shape instead of the real
`{ parallelRouter: "none" }` aggregateDrawElement actually produces.
Extracted applyDeParallelRouterOutcome to keep maybeConsumeDeParallelRouterTrial
under the repo's complexity gate after adding the retry loop.
11 new/updated tests in render.test.ts (56 total) + 7 new tests in
config.test.ts. Verified against fallow's audit gate clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Four confirmed findings from a max-effort code review of the CLI trial
mechanism:
1. maybeEnableDeParallelRouterTrial's `process.env.HF_DE_PARALLEL_ROUTER
!== undefined` guard couldn't distinguish "the user set this" from "an
earlier renderLocal() call in this same process already armed it" — so
in --batch (all rows share one process), only row 1's outcome could
ever reach maybeConsumeDeParallelRouterTrial. A revert on any later row
was silently never persisted. Added a module-level
deParallelRouterTrialManagedByUs flag to disambiguate, with a test-only
reset export since it's process-lifetime state a real CLI invocation
never needs to reset but a test suite sharing one module instance does.
2. writeConfig is a non-atomic whole-file overwrite with no locking, and
readConfig's cache never invalidates — a concurrently running second
CLI process (another terminal, a parallel script; doesn't even need to
be a render, any command calls incrementCommandCount) could silently
clobber a just-persisted deParallelRouterTrialFired:true with its own
stale snapshot. Added readConfigFresh (bypasses the cache) and use it
immediately before the trial's read-modify-write, narrowing the race
window without a full config-subsystem locking rewrite.
3. The prior commit's semantics flip removed the only exposure cap — a
healthy router that never reverts now force-enabled the experimental
path on every eligible render forever. Added
DE_PARALLEL_ROUTER_TRIAL_MAX_RENDERS (25) as a backstop: the trial turns
off after this many engaged renders even absent an actual failure.
4. maybeEnableDeParallelRouterTrial only checked config.telemetryEnabled,
not shouldTrack() — so a dev-mode run or a DO_NOT_TRACK/
HYPERFRAMES_NO_TELEMETRY user got the experimental path silently armed
while telemetry was simultaneously blocked underneath it. Now gates on
shouldTrack() (a strict superset).
Also fixed, lower severity: readConfig's deParallelRouterTrialFired/
deParallelRouterTrialRenderCount parsing now validates the JSON type
explicitly instead of a bare truthy/nullish read, so a hand-edited or
corrupted config can't have the string "false" misread as truthy.
Refactored maybeEnableDeParallelRouterTrial into three smaller functions
(isDeParallelRouterTrialBlocked, stopManagingDeParallelRouterTrial) to
bring cyclomatic/cognitive complexity back under the repo's threshold —
also de-duplicates the "stop managing the env var" logic shared with
maybeConsumeDeParallelRouterTrial.
14 new/updated tests (43 total in render.test.ts), including a direct
regression test for the batch re-entrancy scenario and a loop test for the
render-count cap. Verified the config primitives end-to-end against a real
file, not just the mocked unit tests.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>