mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-03 04:38:33 +00:00
* feat(gcp-cloud-run): add Google Cloud Run + Workflows distributed render adapter Adds @hyperframes/gcp-cloud-run, the GCP counterpart to @hyperframes/aws-lambda (issue #932). The OSS distributed primitives (plan, renderChunk x N, assemble) are unchanged; this package is the storage/compute/orchestration glue. Package: Cloud Run handler (one image, three actions), runs under bun; GCS transport; in-image chrome-headless-shell resolver; client SDK (renderToCloudRun, getRenderProgress, deploySite, computeRenderCost); Dockerfile; Cloud Workflows definition; Terraform module; CLI cloudrun deploy|sites|render|render-batch|progress|destroy with --output-resolution and --strict-variables; 62 unit tests + docs + live smoke script. Shared extraction (removes ~640 lines of adapter duplication): move the cloud-agnostic config validator + content-hash into producer/distributed; both adapters import them. Validated end-to-end on GCP at 37.4 dB PSNR vs baseline. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(cli): resolve @hyperframes/gcp-cloud-run in the CLI build + root build The CLI bundle (esbuild) couldn't resolve `@hyperframes/gcp-cloud-run/sdk`, failing Build/Typecheck/CLI-smoke (and the perf/windows/regression jobs that build first). Mirror the aws-lambda handling: mark the gcp adapter + its /sdk subpath external in tsup.config.ts with a source alias, and add gcp-cloud-run to the root `build` filter so its dist exists for publish + runtime. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ci): copy gcp-cloud-run manifest in Dockerfile.test for frozen install The regression test image runs `bun install --frozen-lockfile` after copying each workspace package.json individually. The CLI now depends on @hyperframes/gcp-cloud-run (workspace:*), so the frozen install fails to resolve it unless its manifest is present. Add the COPY line. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(cli): add machine-sizing flags to `cloudrun deploy` Closes the parity gap with `lambda deploy` (which exposes --memory etc.). `cloudrun deploy` now threads --cpu, --memory, --max-instances, and --timeout into the Terraform apply; omitted flags keep the module defaults (4 vCPU / 16Gi / 100 instances / 3600s). For finer control, apply the module directly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(gcp-cloud-run): address PR review (security, waste, limits, alerts) - server.ts: bucket-allowlist guard no longer fails open silently. Unset env logs a one-time WARNING; "*" is an explicit opt-out; otherwise it enforces. - server.ts: stop double-shipping audio.aac. It already rides in the plan tarball every consumer downloads, so drop the redundant standalone upload (plan) + re-download/overwrite (assemble); assemble reads it from the untar, falling back to a supplied AudioGcsUri for compat. - server.ts: chunk extension via path.extname() instead of slice(lastIndexOf). - workflow.yaml: clamp parallel concurrency_limit to math.min(chunkCount, 20) — Cloud Workflows hard-caps concurrent iterations at 20. - Dockerfile: pin bun (bun-v1.3.9) so an interop change can't silently break the image rebuild. - terraform: add min_instances var (default 0); add a workflow-failure alert (finished_execution_count status=FAILED) alongside the request-count one. - costAccounting: document that displayCost excludes GCS storage/egress. Verified against the actual APIs: @google-cloud/workflows@4.4.0 ICreateExecutionRequest has no executionId (so the idempotency-token suggestion isn't available in this client); Workflows concurrency cap is 20; failure metric is workflows.googleapis.com/finished_execution_count (status label). 174 adapter tests pass, fallow/oxlint/oxfmt/terraform clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(gcp-cloud-run): address round-2 review — error code + CFR forwarding - workflow.yaml: rename the zero-chunk failure code PLAN_TOO_LARGE → PLAN_PRODUCED_ZERO_CHUNKS. The old code implied a size-ceiling breach (the opposite cause), misleading anyone triaging the alert. - workflow.yaml: forward Config.cfr to the assemble step (`Cfr: ${("cfr" in config) and config.cfr}`). It was read by the handler but never sent, so exact-CFR was silently off for every Cloud Run render. Uses the same `in`-operator guard already proven in the retryable predicate. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(release): include gcp-cloud-run in set-version PACKAGES list set-version.ts (driven by release:prepare) bumps an explicit package list to the shared version on each release. gcp-cloud-run was wired into the build + publish.yml but missing here, so a release would leave it at a stale version and publish.yml would push the wrong version. Add it so the new package version-bumps + publishes in lockstep with the others. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
129 lines
5.2 KiB
Markdown
129 lines
5.2 KiB
Markdown
# @hyperframes/gcp-cloud-run
|
||
|
||
Google Cloud Run + Cloud Workflows adapter for HyperFrames distributed
|
||
rendering. The OSS render primitives (`plan` → `renderChunk` × N →
|
||
`assemble`) are pure functions over local file paths; this package is the
|
||
deployment, orchestration, and storage glue that runs them on Google Cloud —
|
||
the GCP counterpart to [`@hyperframes/aws-lambda`](../aws-lambda).
|
||
|
||
Two surfaces, one package:
|
||
|
||
- **Server-side handler** (`./server`) — a Cloud Run HTTP service that
|
||
dispatches `plan` / `renderChunk` / `assemble` on the request body's
|
||
`Action` field, bridging GCS ↔ the container's filesystem around each OSS
|
||
primitive. This is what the bundled `Dockerfile` runs.
|
||
- **Client-side SDK** (`./sdk`) — `renderToCloudRun`, `getRenderProgress`,
|
||
`deploySite`, `validateDistributedRenderConfig`, and `computeRenderCost`.
|
||
Call these from a Node process (CI, CLI, app backend) to drive a deployed
|
||
stack without writing GCS / Workflows boilerplate.
|
||
|
||
The package is **not** a dependency of `@hyperframes/producer`; install it
|
||
separately.
|
||
|
||
## Architecture
|
||
|
||
```
|
||
GCS bucket ←→ Cloud Run service (plan / renderChunk / assemble)
|
||
▲
|
||
│ OIDC-authenticated http.post, one per step
|
||
│
|
||
Cloud Workflows (Plan → parallel RenderChunk → Assemble)
|
||
```
|
||
|
||
- **Plan** downloads the project tarball, runs `plan()`, uploads the planDir
|
||
tarball (+ audio) to GCS, and returns the chunk count.
|
||
- **RenderChunk** runs in a parallel `for` loop in the workflow, fanned out
|
||
up to the plan's chunk count. Each invocation renders one chunk and uploads
|
||
it.
|
||
- **Assemble** downloads every chunk + audio, stitches the final
|
||
deliverable, and uploads it.
|
||
|
||
Every step is a `POST` to the same Cloud Run URL with a different `Action`.
|
||
The workflow accumulates each step's small result body and returns
|
||
`{ Plan, Chunks, Assemble }` so `getRenderProgress` can read frame totals and
|
||
per-step durations on success.
|
||
|
||
## Chrome runtime
|
||
|
||
Unlike the Lambda adapter — which fights a 250 MB ZIP ceiling and
|
||
decompresses `@sparticuz/chromium` into `/tmp` at runtime — Cloud Run runs a
|
||
container image. The `Dockerfile` installs the same pinned
|
||
`chrome-headless-shell` build and font set the production renderer uses, at a
|
||
fixed path, and exports `HYPERFRAMES_CHROME_PATH`. CDP-level `BeginFrame`
|
||
works because the command lives in the protocol, not the binary. There is no
|
||
runtime decompression step and no packaging ceiling.
|
||
|
||
## Deploying
|
||
|
||
The `terraform/` module provisions everything: the GCS render bucket, the
|
||
Cloud Run service, the Cloud Workflows definition, two least-privilege
|
||
service accounts (the service reads/writes the bucket; the workflow invokes
|
||
the service), and a runaway-request alert.
|
||
|
||
```bash
|
||
# 1. Build + push the image (Cloud Build or local docker).
|
||
gcloud builds submit . \
|
||
--tag REGION-docker.pkg.dev/PROJECT/REPO/hyperframes-render:TAG
|
||
|
||
# 2. Apply the module.
|
||
terraform -chdir=node_modules/@hyperframes/gcp-cloud-run/terraform init
|
||
terraform -chdir=node_modules/@hyperframes/gcp-cloud-run/terraform apply \
|
||
-var project_id=PROJECT \
|
||
-var region=us-central1 \
|
||
-var image=REGION-docker.pkg.dev/PROJECT/REPO/hyperframes-render:TAG
|
||
```
|
||
|
||
Terraform outputs `render_bucket_name`, `service_url`, `workflow_name`, and
|
||
`region` — pass them straight into the SDK.
|
||
|
||
## Using the SDK
|
||
|
||
```ts
|
||
import { renderToCloudRun, getRenderProgress } from "@hyperframes/gcp-cloud-run/sdk";
|
||
|
||
const handle = await renderToCloudRun({
|
||
projectDir: "./my-composition",
|
||
config: { fps: 30, width: 1920, height: 1080, format: "mp4" },
|
||
bucketName: "hyperframes-render-my-project", // from terraform output
|
||
projectId: "my-project",
|
||
location: "us-central1",
|
||
workflowId: "hyperframes-render",
|
||
serviceUrl: "https://hyperframes-render-abc.us-central1.run.app",
|
||
});
|
||
|
||
// Poll until done.
|
||
let progress = await getRenderProgress({ executionName: handle.executionName });
|
||
while (progress.status === "running") {
|
||
await new Promise((r) => setTimeout(r, 5000));
|
||
progress = await getRenderProgress({ executionName: handle.executionName });
|
||
}
|
||
console.log(progress.status, progress.outputFile, progress.costs.displayCost);
|
||
```
|
||
|
||
`deploySite` is called implicitly when you pass `projectDir`; call it
|
||
yourself to pre-upload once and reuse the `siteHandle` across many renders
|
||
(e.g. personalised template batches).
|
||
|
||
## Running tests
|
||
|
||
```bash
|
||
bun test # unit tests over an in-memory GCS double — no network
|
||
bun run typecheck
|
||
```
|
||
|
||
The live end-to-end smoke (build image → terraform apply → render a fixture
|
||
through the workflow → PSNR-compare → destroy) lives at
|
||
`examples/gcp-cloud-run/scripts/smoke.sh` and needs a GCP project with
|
||
billing enabled.
|
||
|
||
## What's still ahead
|
||
|
||
- **Mid-flight per-chunk progress.** `getRenderProgress` reports coarse
|
||
`running` progress and exact numbers on success. Reading the Cloud
|
||
Workflows step-entries API would give per-chunk progress while the render
|
||
is in flight; tracked as a follow-up.
|
||
- **Cloud Run Jobs / Firebase Functions variants.** This first version
|
||
targets Cloud Run services + Workflows (the closest analog to Lambda +
|
||
Step Functions). The same handler runs unchanged under Cloud Run Jobs;
|
||
only the orchestration trigger differs.
|