Files
hyperframes/packages/aws-lambda
James RussoandClaude Opus 4.8 4da567df22 feat(gcp-cloud-run): Google Cloud Run + Workflows distributed render adapter (#1253)
* feat(gcp-cloud-run): add Google Cloud Run + Workflows distributed render adapter

Adds @hyperframes/gcp-cloud-run, the GCP counterpart to @hyperframes/aws-lambda
(issue #932). The OSS distributed primitives (plan, renderChunk x N, assemble)
are unchanged; this package is the storage/compute/orchestration glue.

Package: Cloud Run handler (one image, three actions), runs under bun; GCS
transport; in-image chrome-headless-shell resolver; client SDK
(renderToCloudRun, getRenderProgress, deploySite, computeRenderCost); Dockerfile;
Cloud Workflows definition; Terraform module; CLI cloudrun
deploy|sites|render|render-batch|progress|destroy with --output-resolution and
--strict-variables; 62 unit tests + docs + live smoke script.

Shared extraction (removes ~640 lines of adapter duplication): move the
cloud-agnostic config validator + content-hash into producer/distributed; both
adapters import them. Validated end-to-end on GCP at 37.4 dB PSNR vs baseline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(cli): resolve @hyperframes/gcp-cloud-run in the CLI build + root build

The CLI bundle (esbuild) couldn't resolve `@hyperframes/gcp-cloud-run/sdk`,
failing Build/Typecheck/CLI-smoke (and the perf/windows/regression jobs that
build first). Mirror the aws-lambda handling: mark the gcp adapter + its /sdk
subpath external in tsup.config.ts with a source alias, and add gcp-cloud-run
to the root `build` filter so its dist exists for publish + runtime.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ci): copy gcp-cloud-run manifest in Dockerfile.test for frozen install

The regression test image runs `bun install --frozen-lockfile` after copying
each workspace package.json individually. The CLI now depends on
@hyperframes/gcp-cloud-run (workspace:*), so the frozen install fails to
resolve it unless its manifest is present. Add the COPY line.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(cli): add machine-sizing flags to `cloudrun deploy`

Closes the parity gap with `lambda deploy` (which exposes --memory etc.).
`cloudrun deploy` now threads --cpu, --memory, --max-instances, and --timeout
into the Terraform apply; omitted flags keep the module defaults
(4 vCPU / 16Gi / 100 instances / 3600s). For finer control, apply the module
directly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(gcp-cloud-run): address PR review (security, waste, limits, alerts)

- server.ts: bucket-allowlist guard no longer fails open silently. Unset env
  logs a one-time WARNING; "*" is an explicit opt-out; otherwise it enforces.
- server.ts: stop double-shipping audio.aac. It already rides in the plan
  tarball every consumer downloads, so drop the redundant standalone upload
  (plan) + re-download/overwrite (assemble); assemble reads it from the untar,
  falling back to a supplied AudioGcsUri for compat.
- server.ts: chunk extension via path.extname() instead of slice(lastIndexOf).
- workflow.yaml: clamp parallel concurrency_limit to math.min(chunkCount, 20)
  — Cloud Workflows hard-caps concurrent iterations at 20.
- Dockerfile: pin bun (bun-v1.3.9) so an interop change can't silently break
  the image rebuild.
- terraform: add min_instances var (default 0); add a workflow-failure alert
  (finished_execution_count status=FAILED) alongside the request-count one.
- costAccounting: document that displayCost excludes GCS storage/egress.

Verified against the actual APIs: @google-cloud/workflows@4.4.0
ICreateExecutionRequest has no executionId (so the idempotency-token suggestion
isn't available in this client); Workflows concurrency cap is 20; failure
metric is workflows.googleapis.com/finished_execution_count (status label).
174 adapter tests pass, fallow/oxlint/oxfmt/terraform clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(gcp-cloud-run): address round-2 review — error code + CFR forwarding

- workflow.yaml: rename the zero-chunk failure code PLAN_TOO_LARGE →
  PLAN_PRODUCED_ZERO_CHUNKS. The old code implied a size-ceiling breach (the
  opposite cause), misleading anyone triaging the alert.
- workflow.yaml: forward Config.cfr to the assemble step
  (`Cfr: ${("cfr" in config) and config.cfr}`). It was read by the handler
  but never sent, so exact-CFR was silently off for every Cloud Run render.
  Uses the same `in`-operator guard already proven in the retryable predicate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(release): include gcp-cloud-run in set-version PACKAGES list

set-version.ts (driven by release:prepare) bumps an explicit package list to
the shared version on each release. gcp-cloud-run was wired into the build +
publish.yml but missing here, so a release would leave it at a stale version
and publish.yml would push the wrong version. Add it so the new package
version-bumps + publishes in lockstep with the others.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 14:43:38 -07:00
..
2026-06-07 13:35:24 +00:00

@hyperframes/aws-lambda

AWS Lambda adapter for HyperFrames distributed rendering. Ships three things together:

  1. The Lambda handler that wraps the OSS plan / renderChunk / assemble primitives behind a single dispatch boundary Step Functions can drive (src/handler.ts).
  2. A client-side SDKrenderToLambda, getRenderProgress, deploySite, plus validateDistributedRenderConfig and computeRenderCost (src/sdk/).
  3. An aws-cdk-lib L2 construct (HyperframesRenderStack) that provisions the same topology as examples/aws-lambda/template.yaml inside an adopter's own CDK app (src/cdk/).

The handler ZIP and the SAM template still drive a maintainer-run real-AWS smoke flow; the SDK + CDK are the supported public surface for adopters.

Architecture

┌──────────────────────────────────────────────────────────────────┐
│ Step Functions state machine                                     │
│   Plan → Map(N) RenderChunk → Assemble                           │
└──────────────────────────────────────────────────────────────────┘
                              │ dispatches by event.Action
                              ▼
┌──────────────────────────────────────────────────────────────────┐
│ One Lambda function (this package's `dist/handler.zip`)          │
│   handler.mjs                                                    │
│     ├─ Action="plan"        → @hyperframes/producer/distributed  │
│     ├─ Action="renderChunk" → @hyperframes/producer/distributed  │
│     └─ Action="assemble"    → @hyperframes/producer/distributed  │
│   bin/ffmpeg                — ffmpeg-static                      │
│   node_modules/@sparticuz/chromium/ — Lambda-optimised Chromium  │
└──────────────────────────────────────────────────────────────────┘
                              │ pure functions over local paths
                              ▼
┌──────────────────────────────────────────────────────────────────┐
│ S3 bucket — plan tarball + per-chunk outputs + final mp4         │
└──────────────────────────────────────────────────────────────────┘

The handler downloads inputs from S3 into /tmp, calls the OSS primitive, uploads outputs back to S3, and returns a small JSON result that fits inside Step Functions' history budget (under 200 bytes per chunk).

Chrome runtime

The package supports two Chromium sources:

Source Default Size When to pick it
@sparticuz/chromium yes ~70 MiB compressed Lambda. Decompresses into /tmp at runtime; the rest of the ecosystem already uses it for headless-Chrome-in-Lambda.
Bundled chrome-headless-shell no ~140 MiB Fallback. Used if @sparticuz/chromium ever drops HeadlessExperimental.beginFrame support.

Pick the source at build time:

bun run --cwd packages/aws-lambda build:zip
bun run --cwd packages/aws-lambda build:zip -- --source=chrome-headless-shell

The handler reads HYPERFRAMES_LAMBDA_CHROME_SOURCE at boot. The build script sets that env var via Lambda function configuration in examples/aws-lambda/template.yaml.

BeginFrame regression guard

HyperFrames' renderer drives Chrome via the CDP HeadlessExperimental.beginFrame command — same path the K8s deploy uses. The Lambda adapter assumes that @sparticuz/chromium's chrome-headless-shell build honours BeginFrame. To prove it (and re-prove it on every release), the package ships a Docker probe:

# Build the Lambda-like container and run the probe.
bun run --cwd packages/aws-lambda probe:beginframe:docker

The probe boots @sparticuz/chromium inside public.ecr.aws/lambda/nodejs:22 and asserts CDP beginFrame with screenshot: true returns a PNG buffer. Exit code 0 = green; non-zero = fall back to bundling chrome-headless-shell directly via --source=chrome-headless-shell.

Building the ZIP

bun install                                          # at the monorepo root
bun run --cwd packages/aws-lambda build:zip          # → packages/aws-lambda/dist/handler.zip
bun run --cwd packages/aws-lambda verify:zip-size    # CI gate

The build script bundles src/handler.ts via esbuild, stages @sparticuz/chromium and puppeteer-core under node_modules/, copies ffmpeg-static into bin/, and zips the result. The unzipped layout is designed to extract cleanly into Lambda's /var/task/.

verify:zip-size enforces:

  • Unzipped ≤ 248 MiB (in-house budget; Lambda hard ceiling is 250 MiB unzipped — AWS docs label this "250 MB" but use binary mebibytes)
  • Zipped ≤ 150 MiB (in-house budget; Lambda has no hard zipped cap for S3-deployed functions)

CI fails the PR if either is exceeded.

Running tests

bun run --cwd packages/aws-lambda test               # unit tests (no Chrome)
bun run --cwd packages/aws-lambda probe:beginframe   # local probe (Linux only)

Using the SDK

After deploying the stack (via the SAM template, CDK construct below, or your own CFN of choice), drive renders from Node:

import { deploySite, getRenderProgress, renderToLambda } from "@hyperframes/aws-lambda";

// One-time upload per project version.
const site = await deploySite({
  projectDir: "./my-composition",
  bucketName: "hyperframes-render-bucket",
});

// Start a render. Returns immediately — does NOT poll.
const handle = await renderToLambda({
  siteHandle: site,
  bucketName: site.bucketName,
  stateMachineArn: "arn:aws:states:us-east-1:123:stateMachine:hyperframes-render",
  config: {
    fps: 30,
    width: 1920,
    height: 1080,
    format: "mp4",
    chunkSize: 240,
    maxParallelChunks: 16,
    runtimeCap: "lambda",
  },
});

// Poll progress + cost on your own cadence.
const progress = await getRenderProgress({ executionArn: handle.executionArn });
console.log(progress.overallProgress, progress.costs.displayCost);
if (progress.status === "SUCCEEDED" && progress.outputFile) {
  console.log("Render landed at", progress.outputFile.s3Uri);
}

renderToLambda validates the config client-side via validateDistributedRenderConfig and throws a typed InvalidConfigError before the Step Functions execution starts, so shape errors surface synchronously instead of as opaque ExecutionFailed results.

getRenderProgress reports an approximate per-render cost (accruedSoFarUsd plus a formatted displayCost) derived from Lambda billed-duration × memory × the us-east-1 on-demand rate plus the Step Functions transition price. The math is documented in src/sdk/costAccounting.ts; numbers are best-effort and exclude S3 transfer.

Using the CDK construct

import { App, Stack } from "aws-cdk-lib";
import { HyperframesRenderStack } from "@hyperframes/aws-lambda/cdk";

const app = new App();
const stack = new Stack(app, "MyApp");
const render = new HyperframesRenderStack(stack, "Render", {
  // optional: reservedConcurrency: 8,
  // optional: lambdaMemoryMb: 10240,
  // optional: chromeSource: "sparticuz",
});

// Re-export so an adopter app can wire dashboards / SNS topics.
new CfnOutput(stack, "RenderBucketName", { value: render.bucket.bucketName });
new CfnOutput(stack, "StateMachineArn", { value: render.stateMachine.stateMachineArn });

aws-cdk-lib and constructs are optional peer dependencies: SDK-only consumers don't pull them at runtime. The construct itself imports from @hyperframes/aws-lambda/cdk.

What's still ahead

  • hyperframes lambda CLI (deploy / sites create / render / progress / destroy) — PR 6.5.
  • IAM bootstrap subcommand (policies role | user | validate) — PR 6.9.
  • Lambda-local regression harness (--mode=lambda-local) — PR 6.6.
  • Adopter-facing migration guide — PR 6.8.