Files
hyperframes/packages/gcp-cloud-run/README.md
T
James RussoandClaude Opus 4.8 4da567df22 feat(gcp-cloud-run): Google Cloud Run + Workflows distributed render adapter (#1253)
* feat(gcp-cloud-run): add Google Cloud Run + Workflows distributed render adapter

Adds @hyperframes/gcp-cloud-run, the GCP counterpart to @hyperframes/aws-lambda
(issue #932). The OSS distributed primitives (plan, renderChunk x N, assemble)
are unchanged; this package is the storage/compute/orchestration glue.

Package: Cloud Run handler (one image, three actions), runs under bun; GCS
transport; in-image chrome-headless-shell resolver; client SDK
(renderToCloudRun, getRenderProgress, deploySite, computeRenderCost); Dockerfile;
Cloud Workflows definition; Terraform module; CLI cloudrun
deploy|sites|render|render-batch|progress|destroy with --output-resolution and
--strict-variables; 62 unit tests + docs + live smoke script.

Shared extraction (removes ~640 lines of adapter duplication): move the
cloud-agnostic config validator + content-hash into producer/distributed; both
adapters import them. Validated end-to-end on GCP at 37.4 dB PSNR vs baseline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(cli): resolve @hyperframes/gcp-cloud-run in the CLI build + root build

The CLI bundle (esbuild) couldn't resolve `@hyperframes/gcp-cloud-run/sdk`,
failing Build/Typecheck/CLI-smoke (and the perf/windows/regression jobs that
build first). Mirror the aws-lambda handling: mark the gcp adapter + its /sdk
subpath external in tsup.config.ts with a source alias, and add gcp-cloud-run
to the root `build` filter so its dist exists for publish + runtime.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ci): copy gcp-cloud-run manifest in Dockerfile.test for frozen install

The regression test image runs `bun install --frozen-lockfile` after copying
each workspace package.json individually. The CLI now depends on
@hyperframes/gcp-cloud-run (workspace:*), so the frozen install fails to
resolve it unless its manifest is present. Add the COPY line.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(cli): add machine-sizing flags to `cloudrun deploy`

Closes the parity gap with `lambda deploy` (which exposes --memory etc.).
`cloudrun deploy` now threads --cpu, --memory, --max-instances, and --timeout
into the Terraform apply; omitted flags keep the module defaults
(4 vCPU / 16Gi / 100 instances / 3600s). For finer control, apply the module
directly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(gcp-cloud-run): address PR review (security, waste, limits, alerts)

- server.ts: bucket-allowlist guard no longer fails open silently. Unset env
  logs a one-time WARNING; "*" is an explicit opt-out; otherwise it enforces.
- server.ts: stop double-shipping audio.aac. It already rides in the plan
  tarball every consumer downloads, so drop the redundant standalone upload
  (plan) + re-download/overwrite (assemble); assemble reads it from the untar,
  falling back to a supplied AudioGcsUri for compat.
- server.ts: chunk extension via path.extname() instead of slice(lastIndexOf).
- workflow.yaml: clamp parallel concurrency_limit to math.min(chunkCount, 20)
  — Cloud Workflows hard-caps concurrent iterations at 20.
- Dockerfile: pin bun (bun-v1.3.9) so an interop change can't silently break
  the image rebuild.
- terraform: add min_instances var (default 0); add a workflow-failure alert
  (finished_execution_count status=FAILED) alongside the request-count one.
- costAccounting: document that displayCost excludes GCS storage/egress.

Verified against the actual APIs: @google-cloud/workflows@4.4.0
ICreateExecutionRequest has no executionId (so the idempotency-token suggestion
isn't available in this client); Workflows concurrency cap is 20; failure
metric is workflows.googleapis.com/finished_execution_count (status label).
174 adapter tests pass, fallow/oxlint/oxfmt/terraform clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(gcp-cloud-run): address round-2 review — error code + CFR forwarding

- workflow.yaml: rename the zero-chunk failure code PLAN_TOO_LARGE →
  PLAN_PRODUCED_ZERO_CHUNKS. The old code implied a size-ceiling breach (the
  opposite cause), misleading anyone triaging the alert.
- workflow.yaml: forward Config.cfr to the assemble step
  (`Cfr: ${("cfr" in config) and config.cfr}`). It was read by the handler
  but never sent, so exact-CFR was silently off for every Cloud Run render.
  Uses the same `in`-operator guard already proven in the retryable predicate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(release): include gcp-cloud-run in set-version PACKAGES list

set-version.ts (driven by release:prepare) bumps an explicit package list to
the shared version on each release. gcp-cloud-run was wired into the build +
publish.yml but missing here, so a release would leave it at a stale version
and publish.yml would push the wrong version. Add it so the new package
version-bumps + publishes in lockstep with the others.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 14:43:38 -07:00

129 lines
5.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# @hyperframes/gcp-cloud-run
Google Cloud Run + Cloud Workflows adapter for HyperFrames distributed
rendering. The OSS render primitives (`plan``renderChunk` × N →
`assemble`) are pure functions over local file paths; this package is the
deployment, orchestration, and storage glue that runs them on Google Cloud —
the GCP counterpart to [`@hyperframes/aws-lambda`](../aws-lambda).
Two surfaces, one package:
- **Server-side handler** (`./server`) — a Cloud Run HTTP service that
dispatches `plan` / `renderChunk` / `assemble` on the request body's
`Action` field, bridging GCS ↔ the container's filesystem around each OSS
primitive. This is what the bundled `Dockerfile` runs.
- **Client-side SDK** (`./sdk`) — `renderToCloudRun`, `getRenderProgress`,
`deploySite`, `validateDistributedRenderConfig`, and `computeRenderCost`.
Call these from a Node process (CI, CLI, app backend) to drive a deployed
stack without writing GCS / Workflows boilerplate.
The package is **not** a dependency of `@hyperframes/producer`; install it
separately.
## Architecture
```
GCS bucket ←→ Cloud Run service (plan / renderChunk / assemble)
│ OIDC-authenticated http.post, one per step
Cloud Workflows (Plan → parallel RenderChunk → Assemble)
```
- **Plan** downloads the project tarball, runs `plan()`, uploads the planDir
tarball (+ audio) to GCS, and returns the chunk count.
- **RenderChunk** runs in a parallel `for` loop in the workflow, fanned out
up to the plan's chunk count. Each invocation renders one chunk and uploads
it.
- **Assemble** downloads every chunk + audio, stitches the final
deliverable, and uploads it.
Every step is a `POST` to the same Cloud Run URL with a different `Action`.
The workflow accumulates each step's small result body and returns
`{ Plan, Chunks, Assemble }` so `getRenderProgress` can read frame totals and
per-step durations on success.
## Chrome runtime
Unlike the Lambda adapter — which fights a 250 MB ZIP ceiling and
decompresses `@sparticuz/chromium` into `/tmp` at runtime — Cloud Run runs a
container image. The `Dockerfile` installs the same pinned
`chrome-headless-shell` build and font set the production renderer uses, at a
fixed path, and exports `HYPERFRAMES_CHROME_PATH`. CDP-level `BeginFrame`
works because the command lives in the protocol, not the binary. There is no
runtime decompression step and no packaging ceiling.
## Deploying
The `terraform/` module provisions everything: the GCS render bucket, the
Cloud Run service, the Cloud Workflows definition, two least-privilege
service accounts (the service reads/writes the bucket; the workflow invokes
the service), and a runaway-request alert.
```bash
# 1. Build + push the image (Cloud Build or local docker).
gcloud builds submit . \
--tag REGION-docker.pkg.dev/PROJECT/REPO/hyperframes-render:TAG
# 2. Apply the module.
terraform -chdir=node_modules/@hyperframes/gcp-cloud-run/terraform init
terraform -chdir=node_modules/@hyperframes/gcp-cloud-run/terraform apply \
-var project_id=PROJECT \
-var region=us-central1 \
-var image=REGION-docker.pkg.dev/PROJECT/REPO/hyperframes-render:TAG
```
Terraform outputs `render_bucket_name`, `service_url`, `workflow_name`, and
`region` — pass them straight into the SDK.
## Using the SDK
```ts
import { renderToCloudRun, getRenderProgress } from "@hyperframes/gcp-cloud-run/sdk";
const handle = await renderToCloudRun({
projectDir: "./my-composition",
config: { fps: 30, width: 1920, height: 1080, format: "mp4" },
bucketName: "hyperframes-render-my-project", // from terraform output
projectId: "my-project",
location: "us-central1",
workflowId: "hyperframes-render",
serviceUrl: "https://hyperframes-render-abc.us-central1.run.app",
});
// Poll until done.
let progress = await getRenderProgress({ executionName: handle.executionName });
while (progress.status === "running") {
await new Promise((r) => setTimeout(r, 5000));
progress = await getRenderProgress({ executionName: handle.executionName });
}
console.log(progress.status, progress.outputFile, progress.costs.displayCost);
```
`deploySite` is called implicitly when you pass `projectDir`; call it
yourself to pre-upload once and reuse the `siteHandle` across many renders
(e.g. personalised template batches).
## Running tests
```bash
bun test # unit tests over an in-memory GCS double — no network
bun run typecheck
```
The live end-to-end smoke (build image → terraform apply → render a fixture
through the workflow → PSNR-compare → destroy) lives at
`examples/gcp-cloud-run/scripts/smoke.sh` and needs a GCP project with
billing enabled.
## What's still ahead
- **Mid-flight per-chunk progress.** `getRenderProgress` reports coarse
`running` progress and exact numbers on success. Reading the Cloud
Workflows step-entries API would give per-chunk progress while the render
is in flight; tracked as a follow-up.
- **Cloud Run Jobs / Firebase Functions variants.** This first version
targets Cloud Run services + Workflows (the closest analog to Lambda +
Step Functions). The same handler runs unchanged under Cloud Run Jobs;
only the orchestration trigger differs.