mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-04 16:42:27 +00:00
feat(gcp-cloud-run): Google Cloud Run + Workflows distributed render adapter (#1253)
* feat(gcp-cloud-run): add Google Cloud Run + Workflows distributed render adapter Adds @hyperframes/gcp-cloud-run, the GCP counterpart to @hyperframes/aws-lambda (issue #932). The OSS distributed primitives (plan, renderChunk x N, assemble) are unchanged; this package is the storage/compute/orchestration glue. Package: Cloud Run handler (one image, three actions), runs under bun; GCS transport; in-image chrome-headless-shell resolver; client SDK (renderToCloudRun, getRenderProgress, deploySite, computeRenderCost); Dockerfile; Cloud Workflows definition; Terraform module; CLI cloudrun deploy|sites|render|render-batch|progress|destroy with --output-resolution and --strict-variables; 62 unit tests + docs + live smoke script. Shared extraction (removes ~640 lines of adapter duplication): move the cloud-agnostic config validator + content-hash into producer/distributed; both adapters import them. Validated end-to-end on GCP at 37.4 dB PSNR vs baseline. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(cli): resolve @hyperframes/gcp-cloud-run in the CLI build + root build The CLI bundle (esbuild) couldn't resolve `@hyperframes/gcp-cloud-run/sdk`, failing Build/Typecheck/CLI-smoke (and the perf/windows/regression jobs that build first). Mirror the aws-lambda handling: mark the gcp adapter + its /sdk subpath external in tsup.config.ts with a source alias, and add gcp-cloud-run to the root `build` filter so its dist exists for publish + runtime. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ci): copy gcp-cloud-run manifest in Dockerfile.test for frozen install The regression test image runs `bun install --frozen-lockfile` after copying each workspace package.json individually. The CLI now depends on @hyperframes/gcp-cloud-run (workspace:*), so the frozen install fails to resolve it unless its manifest is present. Add the COPY line. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(cli): add machine-sizing flags to `cloudrun deploy` Closes the parity gap with `lambda deploy` (which exposes --memory etc.). `cloudrun deploy` now threads --cpu, --memory, --max-instances, and --timeout into the Terraform apply; omitted flags keep the module defaults (4 vCPU / 16Gi / 100 instances / 3600s). For finer control, apply the module directly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(gcp-cloud-run): address PR review (security, waste, limits, alerts) - server.ts: bucket-allowlist guard no longer fails open silently. Unset env logs a one-time WARNING; "*" is an explicit opt-out; otherwise it enforces. - server.ts: stop double-shipping audio.aac. It already rides in the plan tarball every consumer downloads, so drop the redundant standalone upload (plan) + re-download/overwrite (assemble); assemble reads it from the untar, falling back to a supplied AudioGcsUri for compat. - server.ts: chunk extension via path.extname() instead of slice(lastIndexOf). - workflow.yaml: clamp parallel concurrency_limit to math.min(chunkCount, 20) — Cloud Workflows hard-caps concurrent iterations at 20. - Dockerfile: pin bun (bun-v1.3.9) so an interop change can't silently break the image rebuild. - terraform: add min_instances var (default 0); add a workflow-failure alert (finished_execution_count status=FAILED) alongside the request-count one. - costAccounting: document that displayCost excludes GCS storage/egress. Verified against the actual APIs: @google-cloud/workflows@4.4.0 ICreateExecutionRequest has no executionId (so the idempotency-token suggestion isn't available in this client); Workflows concurrency cap is 20; failure metric is workflows.googleapis.com/finished_execution_count (status label). 174 adapter tests pass, fallow/oxlint/oxfmt/terraform clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(gcp-cloud-run): address round-2 review — error code + CFR forwarding - workflow.yaml: rename the zero-chunk failure code PLAN_TOO_LARGE → PLAN_PRODUCED_ZERO_CHUNKS. The old code implied a size-ceiling breach (the opposite cause), misleading anyone triaging the alert. - workflow.yaml: forward Config.cfr to the assemble step (`Cfr: ${("cfr" in config) and config.cfr}`). It was read by the handler but never sent, so exact-CFR was silently off for every Cloud Run render. Uses the same `in`-operator guard already proven in the retryable predicate. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(release): include gcp-cloud-run in set-version PACKAGES list set-version.ts (driven by release:prepare) bumps an explicit package list to the shared version on each release. gcp-cloud-run was wired into the build + publish.yml but missing here, so a release would leave it at a stale version and publish.yml would push the wrong version. Add it so the new package version-bumps + publishes in lockstep with the others. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
806b226b34
commit
4da567df22
@@ -0,0 +1,179 @@
|
||||
# HyperFrames distributed render orchestration on Cloud Workflows.
|
||||
#
|
||||
# Plan → BuildChunkList → AssertChunkCount → RenderChunks (parallel) → Assemble
|
||||
#
|
||||
# Mirrors the Step Functions state machine in
|
||||
# `examples/aws-lambda/template.yaml`. Every step POSTs to the same Cloud Run
|
||||
# service URL (passed in as `args.ServiceUrl`) and varies only the body's
|
||||
# `Action`. The service returns the step's small result body on 2xx; on a
|
||||
# non-retryable failure it returns HTTP 400, on a retryable failure HTTP 5xx —
|
||||
# the `retryable` predicate below keys off exactly that split.
|
||||
#
|
||||
# The final returned object accumulates every step's result body so
|
||||
# `getRenderProgress` can read frame totals + per-step durations on success:
|
||||
# { Plan: {...}, Chunks: [{...}, ...], Assemble: {...} }
|
||||
#
|
||||
# Deploy with `gcloud workflows deploy` (the Terraform module / the
|
||||
# `hyperframes cloudrun deploy` command do this for you).
|
||||
|
||||
main:
|
||||
params: [args]
|
||||
steps:
|
||||
- init:
|
||||
assign:
|
||||
- serviceUrl: ${args.ServiceUrl}
|
||||
- projectGcsUri: ${args.ProjectGcsUri}
|
||||
- planOutputGcsPrefix: ${args.PlanOutputGcsPrefix}
|
||||
- outputGcsUri: ${args.OutputGcsUri}
|
||||
- config: ${args.Config}
|
||||
|
||||
# ── Plan (Activity A) ────────────────────────────────────────────────────
|
||||
- plan:
|
||||
try:
|
||||
call: http.post
|
||||
args:
|
||||
url: ${serviceUrl}
|
||||
timeout: 1800
|
||||
auth:
|
||||
type: OIDC
|
||||
body:
|
||||
Action: plan
|
||||
ProjectGcsUri: ${projectGcsUri}
|
||||
PlanOutputGcsPrefix: ${planOutputGcsPrefix}
|
||||
Config: ${config}
|
||||
result: planResp
|
||||
retry:
|
||||
predicate: ${retryable}
|
||||
max_retries: 4
|
||||
backoff:
|
||||
initial_delay: 2
|
||||
max_delay: 60
|
||||
multiplier: 2
|
||||
- capturePlan:
|
||||
assign:
|
||||
- planResult: ${planResp.body}
|
||||
- chunkCount: ${planResult.ChunkCount}
|
||||
|
||||
# ── BuildChunkList + AssertChunkCount ──────────────────────────────────────
|
||||
- assertChunkCount:
|
||||
switch:
|
||||
- condition: ${chunkCount > 0}
|
||||
next: buildChunkList
|
||||
next: planProducedZeroChunks
|
||||
- planProducedZeroChunks:
|
||||
raise:
|
||||
code: PLAN_PRODUCED_ZERO_CHUNKS
|
||||
message: "Plan returned ChunkCount=0 — the composition produced no frames. Non-retryable producer-side invariant violation."
|
||||
- buildChunkList:
|
||||
# Pre-size the ordered chunk-URI + per-chunk result lists so the
|
||||
# parallel branches below assign by index (distinct indices, no
|
||||
# read-modify-write race on a shared accumulator).
|
||||
assign:
|
||||
- chunkIndexes: []
|
||||
- chunkUris: []
|
||||
- chunkResults: []
|
||||
- fillLists:
|
||||
for:
|
||||
value: i
|
||||
range: [0, ${chunkCount - 1}]
|
||||
steps:
|
||||
- appendSlots:
|
||||
assign:
|
||||
- chunkIndexes: ${list.concat(chunkIndexes, i)}
|
||||
- chunkUris: ${list.concat(chunkUris, "")}
|
||||
- chunkResults: ${list.concat(chunkResults, "")}
|
||||
|
||||
# ── RenderChunks (Activity B, fanned out) ──────────────────────────────────
|
||||
- renderChunks:
|
||||
parallel:
|
||||
shared: [chunkUris, chunkResults]
|
||||
# Run up to chunkCount chunks at once, clamped to 20 — Cloud
|
||||
# Workflows hard-caps concurrent branches/iterations per execution
|
||||
# at 20 (https://cloud.google.com/workflows/quotas). Above that,
|
||||
# iterations queue regardless of concurrency_limit, so a config with
|
||||
# maxParallelChunks > 20 still renders correctly; the extra chunks
|
||||
# just wait. All chunkCount iterations always run.
|
||||
concurrency_limit: ${math.min(chunkCount, 20)}
|
||||
for:
|
||||
value: idx
|
||||
in: ${chunkIndexes}
|
||||
steps:
|
||||
- renderOneChunk:
|
||||
try:
|
||||
call: http.post
|
||||
args:
|
||||
url: ${serviceUrl}
|
||||
timeout: 1800
|
||||
auth:
|
||||
type: OIDC
|
||||
body:
|
||||
Action: renderChunk
|
||||
ChunkIndex: ${idx}
|
||||
PlanGcsUri: ${planResult.PlanGcsUri}
|
||||
PlanHash: ${planResult.PlanHash}
|
||||
ChunkOutputGcsPrefix: ${planOutputGcsPrefix}
|
||||
Format: ${planResult.Format}
|
||||
result: chunkResp
|
||||
retry:
|
||||
predicate: ${retryable}
|
||||
max_retries: 4
|
||||
backoff:
|
||||
initial_delay: 2
|
||||
max_delay: 60
|
||||
multiplier: 2
|
||||
- storeChunk:
|
||||
assign:
|
||||
- chunkUris[idx]: ${chunkResp.body.ChunkGcsUri}
|
||||
- chunkResults[idx]: ${chunkResp.body}
|
||||
|
||||
# ── Assemble (Activity C) ──────────────────────────────────────────────────
|
||||
- assemble:
|
||||
try:
|
||||
call: http.post
|
||||
args:
|
||||
url: ${serviceUrl}
|
||||
timeout: 1800
|
||||
auth:
|
||||
type: OIDC
|
||||
body:
|
||||
Action: assemble
|
||||
PlanGcsUri: ${planResult.PlanGcsUri}
|
||||
ChunkGcsUris: ${chunkUris}
|
||||
AudioGcsUri: ${planResult.AudioGcsUri}
|
||||
OutputGcsUri: ${outputGcsUri}
|
||||
Format: ${planResult.Format}
|
||||
# Forward the caller's exact-CFR request (Config.cfr) to assemble.
|
||||
# `"cfr" in config` guards the optional key; when unset this is
|
||||
# false, which the handler reads as the default -c copy path.
|
||||
Cfr: ${("cfr" in config) and config.cfr}
|
||||
result: assembleResp
|
||||
retry:
|
||||
predicate: ${retryable}
|
||||
max_retries: 4
|
||||
backoff:
|
||||
initial_delay: 2
|
||||
max_delay: 60
|
||||
multiplier: 2
|
||||
|
||||
- done:
|
||||
return:
|
||||
Plan: ${planResult}
|
||||
Chunks: ${chunkResults}
|
||||
Assemble: ${assembleResp.body}
|
||||
|
||||
# Retry predicate: retry transient/server failures (429 + 5xx), never the
|
||||
# handler's non-retryable 400s (bad input, plan-hash mismatch, unsupported
|
||||
# format, …). Connection / timeout errors carry no `.code`; retry those too.
|
||||
retryable:
|
||||
params: [e]
|
||||
steps:
|
||||
- classify:
|
||||
switch:
|
||||
- condition: ${not("code" in e)}
|
||||
return: true
|
||||
- condition: ${e.code == 429}
|
||||
return: true
|
||||
- condition: ${e.code >= 500 and e.code < 600}
|
||||
return: true
|
||||
- nonRetryable:
|
||||
return: false
|
||||
Reference in New Issue
Block a user