Files
hyperframes/packages/gcp-cloud-run/README.md
T
James RussoandClaude Opus 4.8 4da567df22 feat(gcp-cloud-run): Google Cloud Run + Workflows distributed render adapter (#1253)
* feat(gcp-cloud-run): add Google Cloud Run + Workflows distributed render adapter

Adds @hyperframes/gcp-cloud-run, the GCP counterpart to @hyperframes/aws-lambda
(issue #932). The OSS distributed primitives (plan, renderChunk x N, assemble)
are unchanged; this package is the storage/compute/orchestration glue.

Package: Cloud Run handler (one image, three actions), runs under bun; GCS
transport; in-image chrome-headless-shell resolver; client SDK
(renderToCloudRun, getRenderProgress, deploySite, computeRenderCost); Dockerfile;
Cloud Workflows definition; Terraform module; CLI cloudrun
deploy|sites|render|render-batch|progress|destroy with --output-resolution and
--strict-variables; 62 unit tests + docs + live smoke script.

Shared extraction (removes ~640 lines of adapter duplication): move the
cloud-agnostic config validator + content-hash into producer/distributed; both
adapters import them. Validated end-to-end on GCP at 37.4 dB PSNR vs baseline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(cli): resolve @hyperframes/gcp-cloud-run in the CLI build + root build

The CLI bundle (esbuild) couldn't resolve `@hyperframes/gcp-cloud-run/sdk`,
failing Build/Typecheck/CLI-smoke (and the perf/windows/regression jobs that
build first). Mirror the aws-lambda handling: mark the gcp adapter + its /sdk
subpath external in tsup.config.ts with a source alias, and add gcp-cloud-run
to the root `build` filter so its dist exists for publish + runtime.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ci): copy gcp-cloud-run manifest in Dockerfile.test for frozen install

The regression test image runs `bun install --frozen-lockfile` after copying
each workspace package.json individually. The CLI now depends on
@hyperframes/gcp-cloud-run (workspace:*), so the frozen install fails to
resolve it unless its manifest is present. Add the COPY line.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(cli): add machine-sizing flags to `cloudrun deploy`

Closes the parity gap with `lambda deploy` (which exposes --memory etc.).
`cloudrun deploy` now threads --cpu, --memory, --max-instances, and --timeout
into the Terraform apply; omitted flags keep the module defaults
(4 vCPU / 16Gi / 100 instances / 3600s). For finer control, apply the module
directly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(gcp-cloud-run): address PR review (security, waste, limits, alerts)

- server.ts: bucket-allowlist guard no longer fails open silently. Unset env
  logs a one-time WARNING; "*" is an explicit opt-out; otherwise it enforces.
- server.ts: stop double-shipping audio.aac. It already rides in the plan
  tarball every consumer downloads, so drop the redundant standalone upload
  (plan) + re-download/overwrite (assemble); assemble reads it from the untar,
  falling back to a supplied AudioGcsUri for compat.
- server.ts: chunk extension via path.extname() instead of slice(lastIndexOf).
- workflow.yaml: clamp parallel concurrency_limit to math.min(chunkCount, 20)
  — Cloud Workflows hard-caps concurrent iterations at 20.
- Dockerfile: pin bun (bun-v1.3.9) so an interop change can't silently break
  the image rebuild.
- terraform: add min_instances var (default 0); add a workflow-failure alert
  (finished_execution_count status=FAILED) alongside the request-count one.
- costAccounting: document that displayCost excludes GCS storage/egress.

Verified against the actual APIs: @google-cloud/workflows@4.4.0
ICreateExecutionRequest has no executionId (so the idempotency-token suggestion
isn't available in this client); Workflows concurrency cap is 20; failure
metric is workflows.googleapis.com/finished_execution_count (status label).
174 adapter tests pass, fallow/oxlint/oxfmt/terraform clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(gcp-cloud-run): address round-2 review — error code + CFR forwarding

- workflow.yaml: rename the zero-chunk failure code PLAN_TOO_LARGE →
  PLAN_PRODUCED_ZERO_CHUNKS. The old code implied a size-ceiling breach (the
  opposite cause), misleading anyone triaging the alert.
- workflow.yaml: forward Config.cfr to the assemble step
  (`Cfr: ${("cfr" in config) and config.cfr}`). It was read by the handler
  but never sent, so exact-CFR was silently off for every Cloud Run render.
  Uses the same `in`-operator guard already proven in the retryable predicate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(release): include gcp-cloud-run in set-version PACKAGES list

set-version.ts (driven by release:prepare) bumps an explicit package list to
the shared version on each release. gcp-cloud-run was wired into the build +
publish.yml but missing here, so a release would leave it at a stale version
and publish.yml would push the wrong version. Add it so the new package
version-bumps + publishes in lockstep with the others.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 14:43:38 -07:00

5.2 KiB
Raw Permalink Blame History

@hyperframes/gcp-cloud-run

Google Cloud Run + Cloud Workflows adapter for HyperFrames distributed rendering. The OSS render primitives (planrenderChunk × N → assemble) are pure functions over local file paths; this package is the deployment, orchestration, and storage glue that runs them on Google Cloud — the GCP counterpart to @hyperframes/aws-lambda.

Two surfaces, one package:

  • Server-side handler (./server) — a Cloud Run HTTP service that dispatches plan / renderChunk / assemble on the request body's Action field, bridging GCS ↔ the container's filesystem around each OSS primitive. This is what the bundled Dockerfile runs.
  • Client-side SDK (./sdk) — renderToCloudRun, getRenderProgress, deploySite, validateDistributedRenderConfig, and computeRenderCost. Call these from a Node process (CI, CLI, app backend) to drive a deployed stack without writing GCS / Workflows boilerplate.

The package is not a dependency of @hyperframes/producer; install it separately.

Architecture

GCS bucket  ←→  Cloud Run service (plan / renderChunk / assemble)
                     ▲
                     │ OIDC-authenticated http.post, one per step
                     │
                Cloud Workflows  (Plan → parallel RenderChunk → Assemble)
  • Plan downloads the project tarball, runs plan(), uploads the planDir tarball (+ audio) to GCS, and returns the chunk count.
  • RenderChunk runs in a parallel for loop in the workflow, fanned out up to the plan's chunk count. Each invocation renders one chunk and uploads it.
  • Assemble downloads every chunk + audio, stitches the final deliverable, and uploads it.

Every step is a POST to the same Cloud Run URL with a different Action. The workflow accumulates each step's small result body and returns { Plan, Chunks, Assemble } so getRenderProgress can read frame totals and per-step durations on success.

Chrome runtime

Unlike the Lambda adapter — which fights a 250 MB ZIP ceiling and decompresses @sparticuz/chromium into /tmp at runtime — Cloud Run runs a container image. The Dockerfile installs the same pinned chrome-headless-shell build and font set the production renderer uses, at a fixed path, and exports HYPERFRAMES_CHROME_PATH. CDP-level BeginFrame works because the command lives in the protocol, not the binary. There is no runtime decompression step and no packaging ceiling.

Deploying

The terraform/ module provisions everything: the GCS render bucket, the Cloud Run service, the Cloud Workflows definition, two least-privilege service accounts (the service reads/writes the bucket; the workflow invokes the service), and a runaway-request alert.

# 1. Build + push the image (Cloud Build or local docker).
gcloud builds submit . \
  --tag REGION-docker.pkg.dev/PROJECT/REPO/hyperframes-render:TAG

# 2. Apply the module.
terraform -chdir=node_modules/@hyperframes/gcp-cloud-run/terraform init
terraform -chdir=node_modules/@hyperframes/gcp-cloud-run/terraform apply \
  -var project_id=PROJECT \
  -var region=us-central1 \
  -var image=REGION-docker.pkg.dev/PROJECT/REPO/hyperframes-render:TAG

Terraform outputs render_bucket_name, service_url, workflow_name, and region — pass them straight into the SDK.

Using the SDK

import { renderToCloudRun, getRenderProgress } from "@hyperframes/gcp-cloud-run/sdk";

const handle = await renderToCloudRun({
  projectDir: "./my-composition",
  config: { fps: 30, width: 1920, height: 1080, format: "mp4" },
  bucketName: "hyperframes-render-my-project", // from terraform output
  projectId: "my-project",
  location: "us-central1",
  workflowId: "hyperframes-render",
  serviceUrl: "https://hyperframes-render-abc.us-central1.run.app",
});

// Poll until done.
let progress = await getRenderProgress({ executionName: handle.executionName });
while (progress.status === "running") {
  await new Promise((r) => setTimeout(r, 5000));
  progress = await getRenderProgress({ executionName: handle.executionName });
}
console.log(progress.status, progress.outputFile, progress.costs.displayCost);

deploySite is called implicitly when you pass projectDir; call it yourself to pre-upload once and reuse the siteHandle across many renders (e.g. personalised template batches).

Running tests

bun test          # unit tests over an in-memory GCS double — no network
bun run typecheck

The live end-to-end smoke (build image → terraform apply → render a fixture through the workflow → PSNR-compare → destroy) lives at examples/gcp-cloud-run/scripts/smoke.sh and needs a GCP project with billing enabled.

What's still ahead

  • Mid-flight per-chunk progress. getRenderProgress reports coarse running progress and exact numbers on success. Reading the Cloud Workflows step-entries API would give per-chunk progress while the render is in flight; tracked as a follow-up.
  • Cloud Run Jobs / Firebase Functions variants. This first version targets Cloud Run services + Workflows (the closest analog to Lambda + Step Functions). The same handler runs unchanged under Cloud Run Jobs; only the orchestration trigger differs.