Files
hyperframes/packages/gcp-cloud-run
James Russo 557d82b6a9 fix(producer): validate distributed video metadata (#2839)
## What

- enforce a finite, validated `meta/videos.json` contract shared by Plan v1 and Plan v2
- preserve authored finite ends and source-derived trim-aware ends; bound any still-open end at the validated composition end
- fail distributed planning when any declared video source did not extract instead of publishing a blank-capable plan
- make the v1 chunk reader reject malformed/null video timing before frame injection
- route deterministic video-source/metadata failures as non-retryable in AWS and GCP while retaining retries for transient extraction failures

## Why

An open-ended video whose remote source could not be resolved retained `Infinity` through planning. Plan v2 correctly rejected that value, while Plan v1 serialized it as `null`; the v1 frame lookup could then suppress injected frames and silently produce incorrect output.

The invariant belongs at the shared metadata boundary. Both protocols must receive identical finite timing, and unavailable sources must fail closed before plan publication.

## Test plan

- [x] producer distributed planning, metadata, v1 chunk boundary, Plan v2 conversion/materialization, and public exports
- [x] core runtime media semantics (authored slots, natural duration, looping, non-looping hold)
- [x] engine video extraction and frame lookup
- [x] AWS Lambda/CDK/SAM and GCP Cloud Run error normalization/retry classification
- [x] producer, core, engine, AWS, and GCP typechecks/builds
- [x] formatting, oxlint, tracked-artifact, fallow, and commit hooks
- [x] exact incident composition replayed through the AWS Lambda handler's Lambda-local path in a Lambda-like container; Plan v1 and Plan v2 both fail closed as `VIDEO_SOURCE_UNRENDERABLE` during planning, before plan publication
- [x] full PR CI, including all nine regression shards and Windows render/tests

No production flags or deployment/release workflows are changed.
2026-07-28 00:42:36 -07:00
..
2026-07-28 02:12:41 +00:00

@hyperframes/gcp-cloud-run

Google Cloud Run + Cloud Workflows adapter for HyperFrames distributed rendering. The OSS render primitives (planrenderChunk × N → assemble) are pure functions over local file paths; this package is the deployment, orchestration, and storage glue that runs them on Google Cloud — the GCP counterpart to @hyperframes/aws-lambda.

Two surfaces, one package:

  • Server-side handler (./server) — a Cloud Run HTTP service that dispatches plan / renderChunk / assemble on the request body's Action field, bridging GCS ↔ the container's filesystem around each OSS primitive. This is what the bundled Dockerfile runs.
  • Client-side SDK (./sdk) — renderToCloudRun, getRenderProgress, deploySite, validateDistributedRenderConfig, and computeRenderCost. Call these from a Node process (CI, CLI, app backend) to drive a deployed stack without writing GCS / Workflows boilerplate.

The package is not a dependency of @hyperframes/producer; install it separately.

Architecture

GCS bucket  ←→  Cloud Run service (plan / renderChunk / assemble)
                     ▲
                     │ OIDC-authenticated http.post, one per step
                     │
                Cloud Workflows  (Plan → parallel RenderChunk → Assemble)
  • Plan downloads the project tarball, runs plan(), uploads the planDir tarball (+ audio) to GCS, and returns the chunk count.
  • RenderChunk runs in a parallel for loop in the workflow, fanned out up to the plan's chunk count. Each invocation renders one chunk and uploads it.
  • Assemble downloads every chunk + audio, stitches the final deliverable, and uploads it.

Every step is a POST to the same Cloud Run URL with a different Action. The workflow accumulates each step's small result body and returns { Plan, Chunks, Assemble } so getRenderProgress can read frame totals and per-step durations on success.

Chrome runtime

Unlike the Lambda adapter — which fights a 250 MB ZIP ceiling and decompresses @sparticuz/chromium into /tmp at runtime — Cloud Run runs a container image. The Dockerfile installs the same pinned chrome-headless-shell build and font set the production renderer uses, at a fixed path, and exports HYPERFRAMES_CHROME_PATH. CDP-level BeginFrame support is a binary/runtime capability, so the image build launches that exact executable and requires an enable + warm-up + PNG-returning HeadlessExperimental.beginFrame probe to pass. The end-to-end smoke also requires every chunk to report effective CaptureMode: "beginframe", which catches runtime fallback separately from build-time packaging. There is no runtime decompression step and no packaging ceiling.

Deploying

The terraform/ module provisions everything: the GCS render bucket, the Cloud Run service, the Cloud Workflows definition, two least-privilege service accounts (the service reads/writes the bucket; the workflow invokes the service), and a runaway-request alert.

# 1. Build + push the image (Cloud Build or local docker).
gcloud builds submit . \
  --tag REGION-docker.pkg.dev/PROJECT/REPO/hyperframes-render:TAG

# 2. Apply the module.
terraform -chdir=node_modules/@hyperframes/gcp-cloud-run/terraform init
terraform -chdir=node_modules/@hyperframes/gcp-cloud-run/terraform apply \
  -var project_id=PROJECT \
  -var region=us-central1 \
  -var image=REGION-docker.pkg.dev/PROJECT/REPO/hyperframes-render:TAG

Terraform outputs render_bucket_name, service_url, workflow_name, and region — pass them straight into the SDK.

Using the SDK

import { renderToCloudRun, getRenderProgress } from "@hyperframes/gcp-cloud-run/sdk";

const handle = await renderToCloudRun({
  projectDir: "./my-composition",
  config: { fps: 30, width: 1920, height: 1080, format: "mp4" },
  bucketName: "hyperframes-render-my-project", // from terraform output
  projectId: "my-project",
  location: "us-central1",
  workflowId: "hyperframes-render",
  serviceUrl: "https://hyperframes-render-abc.us-central1.run.app",
});

// Poll until done.
let progress = await getRenderProgress({ executionName: handle.executionName });
while (progress.status === "running") {
  await new Promise((r) => setTimeout(r, 5000));
  progress = await getRenderProgress({ executionName: handle.executionName });
}
console.log(progress.status, progress.outputFile, progress.costs.displayCost);

deploySite is called implicitly when you pass projectDir; call it yourself to pre-upload once and reuse the siteHandle across many renders (e.g. personalised template batches).

Running tests

bun test          # unit tests over an in-memory GCS double — no network
bun run typecheck

The live end-to-end smoke (build image → terraform apply → render a fixture through the workflow → PSNR-compare → destroy) lives at examples/gcp-cloud-run/scripts/smoke.sh and needs a GCP project with billing enabled.

What's still ahead

  • Mid-flight per-chunk progress. getRenderProgress reports coarse running progress and exact numbers on success. Reading the Cloud Workflows step-entries API would give per-chunk progress while the render is in flight; tracked as a follow-up.
  • Cloud Run Jobs / Firebase Functions variants. This first version targets Cloud Run services + Workflows (the closest analog to Lambda + Step Functions). The same handler runs unchanged under Cloud Run Jobs; only the orchestration trigger differs.