mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-03 04:38:33 +00:00
* feat(gcp-cloud-run): add Google Cloud Run + Workflows distributed render adapter Adds @hyperframes/gcp-cloud-run, the GCP counterpart to @hyperframes/aws-lambda (issue #932). The OSS distributed primitives (plan, renderChunk x N, assemble) are unchanged; this package is the storage/compute/orchestration glue. Package: Cloud Run handler (one image, three actions), runs under bun; GCS transport; in-image chrome-headless-shell resolver; client SDK (renderToCloudRun, getRenderProgress, deploySite, computeRenderCost); Dockerfile; Cloud Workflows definition; Terraform module; CLI cloudrun deploy|sites|render|render-batch|progress|destroy with --output-resolution and --strict-variables; 62 unit tests + docs + live smoke script. Shared extraction (removes ~640 lines of adapter duplication): move the cloud-agnostic config validator + content-hash into producer/distributed; both adapters import them. Validated end-to-end on GCP at 37.4 dB PSNR vs baseline. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(cli): resolve @hyperframes/gcp-cloud-run in the CLI build + root build The CLI bundle (esbuild) couldn't resolve `@hyperframes/gcp-cloud-run/sdk`, failing Build/Typecheck/CLI-smoke (and the perf/windows/regression jobs that build first). Mirror the aws-lambda handling: mark the gcp adapter + its /sdk subpath external in tsup.config.ts with a source alias, and add gcp-cloud-run to the root `build` filter so its dist exists for publish + runtime. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ci): copy gcp-cloud-run manifest in Dockerfile.test for frozen install The regression test image runs `bun install --frozen-lockfile` after copying each workspace package.json individually. The CLI now depends on @hyperframes/gcp-cloud-run (workspace:*), so the frozen install fails to resolve it unless its manifest is present. Add the COPY line. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(cli): add machine-sizing flags to `cloudrun deploy` Closes the parity gap with `lambda deploy` (which exposes --memory etc.). `cloudrun deploy` now threads --cpu, --memory, --max-instances, and --timeout into the Terraform apply; omitted flags keep the module defaults (4 vCPU / 16Gi / 100 instances / 3600s). For finer control, apply the module directly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(gcp-cloud-run): address PR review (security, waste, limits, alerts) - server.ts: bucket-allowlist guard no longer fails open silently. Unset env logs a one-time WARNING; "*" is an explicit opt-out; otherwise it enforces. - server.ts: stop double-shipping audio.aac. It already rides in the plan tarball every consumer downloads, so drop the redundant standalone upload (plan) + re-download/overwrite (assemble); assemble reads it from the untar, falling back to a supplied AudioGcsUri for compat. - server.ts: chunk extension via path.extname() instead of slice(lastIndexOf). - workflow.yaml: clamp parallel concurrency_limit to math.min(chunkCount, 20) — Cloud Workflows hard-caps concurrent iterations at 20. - Dockerfile: pin bun (bun-v1.3.9) so an interop change can't silently break the image rebuild. - terraform: add min_instances var (default 0); add a workflow-failure alert (finished_execution_count status=FAILED) alongside the request-count one. - costAccounting: document that displayCost excludes GCS storage/egress. Verified against the actual APIs: @google-cloud/workflows@4.4.0 ICreateExecutionRequest has no executionId (so the idempotency-token suggestion isn't available in this client); Workflows concurrency cap is 20; failure metric is workflows.googleapis.com/finished_execution_count (status label). 174 adapter tests pass, fallow/oxlint/oxfmt/terraform clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(gcp-cloud-run): address round-2 review — error code + CFR forwarding - workflow.yaml: rename the zero-chunk failure code PLAN_TOO_LARGE → PLAN_PRODUCED_ZERO_CHUNKS. The old code implied a size-ceiling breach (the opposite cause), misleading anyone triaging the alert. - workflow.yaml: forward Config.cfr to the assemble step (`Cfr: ${("cfr" in config) and config.cfr}`). It was read by the handler but never sent, so exact-CFR was silently off for every Cloud Run render. Uses the same `in`-operator guard already proven in the retryable predicate. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(release): include gcp-cloud-run in set-version PACKAGES list set-version.ts (driven by release:prepare) bumps an explicit package list to the shared version on each release. gcp-cloud-run was wired into the build + publish.yml but missing here, so a release would leave it at a stale version and publish.yml would push the wrong version. Add it so the new package version-bumps + publishes in lockstep with the others. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
198 lines
7.3 KiB
Terraform
198 lines
7.3 KiB
Terraform
# HyperFrames distributed render stack on Google Cloud.
|
|
#
|
|
# Topology (the GCP twin of the AWS Lambda adapter's SAM template):
|
|
#
|
|
# GCS bucket ←→ Cloud Run service (plan / renderChunk / assemble)
|
|
# ▲
|
|
# │ OIDC-authenticated http.post per step
|
|
# │
|
|
# Cloud Workflows (Plan → parallel RenderChunk → Assemble)
|
|
#
|
|
# Two service accounts keep least-privilege boundaries:
|
|
# - run_sa : the render service's identity; read/write the bucket only.
|
|
# - workflow_sa: the workflow's identity; invoke the render service only.
|
|
|
|
locals {
|
|
workflow_source = var.workflow_source_path != "" ? var.workflow_source_path : "${path.module}/workflow.yaml"
|
|
name = var.project_name
|
|
}
|
|
|
|
# ── Storage: plan tarballs, chunk outputs, final renders ─────────────────────
|
|
resource "google_storage_bucket" "render" {
|
|
name = "${local.name}-render-${var.project_id}"
|
|
project = var.project_id
|
|
location = var.region
|
|
uniform_bucket_level_access = true
|
|
force_destroy = var.bucket_force_destroy
|
|
|
|
# Render artifacts (plan tarballs, chunk files) are disposable scratch.
|
|
# Sweep them after 7 days so the bucket doesn't accumulate cost; final
|
|
# outputs that adopters want to keep should be copied elsewhere.
|
|
lifecycle_rule {
|
|
condition {
|
|
age = 7
|
|
}
|
|
action {
|
|
type = "Delete"
|
|
}
|
|
}
|
|
}
|
|
|
|
# ── Service accounts ─────────────────────────────────────────────────────────
|
|
resource "google_service_account" "run_sa" {
|
|
account_id = "${local.name}-run"
|
|
project = var.project_id
|
|
display_name = "HyperFrames render service (Cloud Run)"
|
|
}
|
|
|
|
resource "google_service_account" "workflow_sa" {
|
|
account_id = "${local.name}-wf"
|
|
project = var.project_id
|
|
display_name = "HyperFrames render orchestration (Workflows)"
|
|
}
|
|
|
|
# Render service reads inputs + writes outputs in the render bucket only.
|
|
resource "google_storage_bucket_iam_member" "run_sa_bucket" {
|
|
bucket = google_storage_bucket.render.name
|
|
role = "roles/storage.objectAdmin"
|
|
member = "serviceAccount:${google_service_account.run_sa.email}"
|
|
}
|
|
|
|
# ── Cloud Run render service ─────────────────────────────────────────────────
|
|
resource "google_cloud_run_v2_service" "render" {
|
|
name = "${local.name}-render"
|
|
project = var.project_id
|
|
location = var.region
|
|
# Authenticated only — no public invoker binding. Only the workflow SA can
|
|
# call it. Ingress stays "all" because Workflows reaches the service over
|
|
# Google's front door, not the VPC.
|
|
ingress = "INGRESS_TRAFFIC_ALL"
|
|
# Let `terraform destroy` (and replacement on image bumps) remove the
|
|
# service without a manual console step. The render service is stateless —
|
|
# all durable artifacts live in GCS.
|
|
deletion_protection = false
|
|
|
|
template {
|
|
service_account = google_service_account.run_sa.email
|
|
timeout = "${var.request_timeout_seconds}s"
|
|
# One render (chunk / plan / assemble) per instance — each uses the whole
|
|
# box's CPU + memory + /tmp. The workflow's concurrency_limit governs how
|
|
# many instances run at once.
|
|
max_instance_request_concurrency = 1
|
|
|
|
scaling {
|
|
min_instance_count = var.min_instances
|
|
max_instance_count = var.max_instances
|
|
}
|
|
|
|
containers {
|
|
image = var.image
|
|
|
|
resources {
|
|
limits = {
|
|
cpu = var.cpu
|
|
memory = var.memory
|
|
}
|
|
# Keep CPU allocated only during request processing (request-based
|
|
# billing). Renders are entirely request-scoped.
|
|
cpu_idle = true
|
|
}
|
|
|
|
env {
|
|
# Scopes every event's GCS URIs to this bucket (the handler's
|
|
# GCS_URI_NOT_ALLOWED guard). Defense against request injection.
|
|
name = "HYPERFRAMES_RENDER_BUCKET"
|
|
value = google_storage_bucket.render.name
|
|
}
|
|
|
|
startup_probe {
|
|
http_get {
|
|
path = "/healthz"
|
|
}
|
|
timeout_seconds = 5
|
|
period_seconds = 10
|
|
failure_threshold = 6
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
# Only the workflow's identity may invoke the render service.
|
|
resource "google_cloud_run_v2_service_iam_member" "workflow_invokes_run" {
|
|
name = google_cloud_run_v2_service.render.name
|
|
project = var.project_id
|
|
location = var.region
|
|
role = "roles/run.invoker"
|
|
member = "serviceAccount:${google_service_account.workflow_sa.email}"
|
|
}
|
|
|
|
# ── Cloud Workflows orchestration ────────────────────────────────────────────
|
|
resource "google_workflows_workflow" "render" {
|
|
name = "${local.name}-render"
|
|
project = var.project_id
|
|
region = var.region
|
|
service_account = google_service_account.workflow_sa.id
|
|
source_contents = file(local.workflow_source)
|
|
# Allow `terraform destroy` to remove the workflow without a manual step;
|
|
# the definition is reproducible from this module.
|
|
deletion_protection = false
|
|
}
|
|
|
|
# ── Runaway-request alert (backstop against a fan-out bug) ────────────────────
|
|
resource "google_monitoring_alert_policy" "runaway_requests" {
|
|
project = var.project_id
|
|
display_name = "${local.name}-render runaway request count"
|
|
combiner = "OR"
|
|
|
|
conditions {
|
|
display_name = "Render service request count > threshold (1h)"
|
|
condition_threshold {
|
|
filter = join(" AND ", [
|
|
"resource.type = \"cloud_run_revision\"",
|
|
"resource.labels.service_name = \"${google_cloud_run_v2_service.render.name}\"",
|
|
"metric.type = \"run.googleapis.com/request_count\"",
|
|
])
|
|
comparison = "COMPARISON_GT"
|
|
threshold_value = var.render_request_alarm_threshold
|
|
duration = "0s"
|
|
aggregations {
|
|
alignment_period = "3600s"
|
|
per_series_aligner = "ALIGN_SUM"
|
|
}
|
|
}
|
|
}
|
|
|
|
notification_channels = var.notification_channels
|
|
}
|
|
|
|
# ── Workflow-failure alert ───────────────────────────────────────────────────
|
|
# Request-count alone misses a render that fails 100% of the time at low
|
|
# volume. Alert on any FAILED workflow execution so a broken render path is
|
|
# visible even when traffic is light.
|
|
resource "google_monitoring_alert_policy" "workflow_failures" {
|
|
project = var.project_id
|
|
display_name = "${local.name}-render workflow execution failures"
|
|
combiner = "OR"
|
|
|
|
conditions {
|
|
display_name = "Failed workflow executions (5m)"
|
|
condition_threshold {
|
|
filter = join(" AND ", [
|
|
"resource.type = \"workflows.googleapis.com/Workflow\"",
|
|
"resource.labels.workflow_id = \"${google_workflows_workflow.render.name}\"",
|
|
"metric.type = \"workflows.googleapis.com/finished_execution_count\"",
|
|
"metric.labels.status = \"FAILED\"",
|
|
])
|
|
comparison = "COMPARISON_GT"
|
|
threshold_value = 0
|
|
duration = "0s"
|
|
aggregations {
|
|
alignment_period = "300s"
|
|
per_series_aligner = "ALIGN_SUM"
|
|
}
|
|
}
|
|
}
|
|
|
|
notification_channels = var.notification_channels
|
|
}
|