mirror of
https://github.com/heygen-com/hyperframes.git
synced 2026-09-05 17:30:50 +00:00
* feat(gcp-cloud-run): add Google Cloud Run + Workflows distributed render adapter Adds @hyperframes/gcp-cloud-run, the GCP counterpart to @hyperframes/aws-lambda (issue #932). The OSS distributed primitives (plan, renderChunk x N, assemble) are unchanged; this package is the storage/compute/orchestration glue. Package: Cloud Run handler (one image, three actions), runs under bun; GCS transport; in-image chrome-headless-shell resolver; client SDK (renderToCloudRun, getRenderProgress, deploySite, computeRenderCost); Dockerfile; Cloud Workflows definition; Terraform module; CLI cloudrun deploy|sites|render|render-batch|progress|destroy with --output-resolution and --strict-variables; 62 unit tests + docs + live smoke script. Shared extraction (removes ~640 lines of adapter duplication): move the cloud-agnostic config validator + content-hash into producer/distributed; both adapters import them. Validated end-to-end on GCP at 37.4 dB PSNR vs baseline. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(cli): resolve @hyperframes/gcp-cloud-run in the CLI build + root build The CLI bundle (esbuild) couldn't resolve `@hyperframes/gcp-cloud-run/sdk`, failing Build/Typecheck/CLI-smoke (and the perf/windows/regression jobs that build first). Mirror the aws-lambda handling: mark the gcp adapter + its /sdk subpath external in tsup.config.ts with a source alias, and add gcp-cloud-run to the root `build` filter so its dist exists for publish + runtime. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ci): copy gcp-cloud-run manifest in Dockerfile.test for frozen install The regression test image runs `bun install --frozen-lockfile` after copying each workspace package.json individually. The CLI now depends on @hyperframes/gcp-cloud-run (workspace:*), so the frozen install fails to resolve it unless its manifest is present. Add the COPY line. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(cli): add machine-sizing flags to `cloudrun deploy` Closes the parity gap with `lambda deploy` (which exposes --memory etc.). `cloudrun deploy` now threads --cpu, --memory, --max-instances, and --timeout into the Terraform apply; omitted flags keep the module defaults (4 vCPU / 16Gi / 100 instances / 3600s). For finer control, apply the module directly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(gcp-cloud-run): address PR review (security, waste, limits, alerts) - server.ts: bucket-allowlist guard no longer fails open silently. Unset env logs a one-time WARNING; "*" is an explicit opt-out; otherwise it enforces. - server.ts: stop double-shipping audio.aac. It already rides in the plan tarball every consumer downloads, so drop the redundant standalone upload (plan) + re-download/overwrite (assemble); assemble reads it from the untar, falling back to a supplied AudioGcsUri for compat. - server.ts: chunk extension via path.extname() instead of slice(lastIndexOf). - workflow.yaml: clamp parallel concurrency_limit to math.min(chunkCount, 20) — Cloud Workflows hard-caps concurrent iterations at 20. - Dockerfile: pin bun (bun-v1.3.9) so an interop change can't silently break the image rebuild. - terraform: add min_instances var (default 0); add a workflow-failure alert (finished_execution_count status=FAILED) alongside the request-count one. - costAccounting: document that displayCost excludes GCS storage/egress. Verified against the actual APIs: @google-cloud/workflows@4.4.0 ICreateExecutionRequest has no executionId (so the idempotency-token suggestion isn't available in this client); Workflows concurrency cap is 20; failure metric is workflows.googleapis.com/finished_execution_count (status label). 174 adapter tests pass, fallow/oxlint/oxfmt/terraform clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(gcp-cloud-run): address round-2 review — error code + CFR forwarding - workflow.yaml: rename the zero-chunk failure code PLAN_TOO_LARGE → PLAN_PRODUCED_ZERO_CHUNKS. The old code implied a size-ceiling breach (the opposite cause), misleading anyone triaging the alert. - workflow.yaml: forward Config.cfr to the assemble step (`Cfr: ${("cfr" in config) and config.cfr}`). It was read by the handler but never sent, so exact-CFR was silently off for every Cloud Run render. Uses the same `in`-operator guard already proven in the retryable predicate. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(release): include gcp-cloud-run in set-version PACKAGES list set-version.ts (driven by release:prepare) bumps an explicit package list to the shared version on each release. gcp-cloud-run was wired into the build + publish.yml but missing here, so a release would leave it at a stale version and publish.yml would push the wrong version. Add it so the new package version-bumps + publishes in lockstep with the others. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
137 lines
5.4 KiB
TypeScript
137 lines
5.4 KiB
TypeScript
/**
|
|
* `deploySite` — upload a project directory to S3 once per content hash
|
|
* and return a reusable handle.
|
|
*
|
|
* `renderToLambda` calls this implicitly when no `siteHandle` is passed,
|
|
* but exposing it as a standalone verb lets adopters bundle a project
|
|
* ahead of time and reuse the handle across many renders without
|
|
* re-tarring the project tree on every call.
|
|
*
|
|
* The handle is **content-addressed**: `siteId` is derived from a SHA-256
|
|
* over the project files. Two `deploySite` calls on an unchanged tree
|
|
* produce the same `siteId` and `HeadObject`-short-circuit the upload.
|
|
*/
|
|
|
|
import { mkdtempSync, rmSync, statSync } from "node:fs";
|
|
import { tmpdir } from "node:os";
|
|
import { join } from "node:path";
|
|
import { HeadObjectCommand, S3Client } from "@aws-sdk/client-s3";
|
|
import { hashProjectDir } from "@hyperframes/producer/distributed";
|
|
import { formatS3Uri, tarDirectory, uploadFileToS3 } from "../s3Transport.js";
|
|
|
|
/** Options for {@link deploySite}. */
|
|
export interface DeploySiteOptions {
|
|
/** Local project directory containing `index.html` (and any composition assets). */
|
|
projectDir: string;
|
|
/** S3 bucket the SAM stack / CDK construct provisioned. */
|
|
bucketName: string;
|
|
/** AWS region for the S3 client. Defaults to the SDK's default chain (env / config / IMDS). */
|
|
region?: string;
|
|
/**
|
|
* Override the content-addressed site id. Useful when the caller has a
|
|
* stable external identifier they want to use (e.g. a git SHA); if
|
|
* unset, the hash of the project tree picks it.
|
|
*/
|
|
siteId?: string;
|
|
/** Injection seam for tests. Production callers leave unset. */
|
|
s3?: S3Client;
|
|
}
|
|
|
|
/** Stable handle returned by {@link deploySite}. Pass back to {@link renderToLambda}. */
|
|
export interface SiteHandle {
|
|
/** Content-addressed (or caller-supplied) identifier; stable across re-uploads of the same tree. */
|
|
siteId: string;
|
|
/** Bucket the site landed in. Surfaced separately so callers don't have to re-parse `projectS3Uri`. */
|
|
bucketName: string;
|
|
/** Full `s3://bucket/sites/<siteId>/project.tar.gz` URI; pass through to `renderToLambda`. */
|
|
projectS3Uri: string;
|
|
/** Tarball size in bytes; useful for "did we actually skip the upload?" assertions. */
|
|
bytes: number;
|
|
/** ISO timestamp of the most recent upload OR the existing object the short-circuit found. */
|
|
uploadedAt: string;
|
|
/** `false` if the object already existed and we skipped the PUT. */
|
|
uploaded: boolean;
|
|
}
|
|
|
|
/**
|
|
* Upload `projectDir` to `s3://bucketName/sites/<siteId>/project.tar.gz`.
|
|
*
|
|
* Short-circuits when an object with the same key already exists in the
|
|
* bucket — `siteId` derives from the project's content hash, so the same
|
|
* bytes produce the same key, and re-uploading would be redundant.
|
|
*/
|
|
export async function deploySite(opts: DeploySiteOptions): Promise<SiteHandle> {
|
|
if (!statSync(opts.projectDir).isDirectory()) {
|
|
throw new Error(`[deploySite] projectDir is not a directory: ${opts.projectDir}`);
|
|
}
|
|
|
|
const siteId = opts.siteId ?? hashProjectDir(opts.projectDir);
|
|
const key = `sites/${siteId}/project.tar.gz`;
|
|
const projectS3Uri = formatS3Uri({ bucket: opts.bucketName, key });
|
|
const s3 = opts.s3 ?? new S3Client({ region: opts.region });
|
|
|
|
// HeadObject short-circuit. Adopters re-rendering the same project on
|
|
// a tight inner loop (CI smoke, demo flows) save the tar+gzip+PUT pass
|
|
// on every iteration.
|
|
const existing = await headObject(s3, opts.bucketName, key);
|
|
if (existing) {
|
|
return {
|
|
siteId,
|
|
bucketName: opts.bucketName,
|
|
projectS3Uri,
|
|
bytes: existing.bytes,
|
|
uploadedAt: existing.lastModified,
|
|
uploaded: false,
|
|
};
|
|
}
|
|
|
|
const workdir = mkdtempSync(join(tmpdir(), "hf-deploy-site-"));
|
|
try {
|
|
const tarball = join(workdir, "project.tar.gz");
|
|
await tarDirectory(opts.projectDir, tarball);
|
|
// Note: tarDirectory packs *everything* under `cwd`. We don't need to
|
|
// re-implement the skip list inside the tar pack because the
|
|
// producer's plan stage applies the same skip during its copy; the
|
|
// archive is slightly bigger than the planDir's compiled/ subtree
|
|
// but the cost is bounded by the project's user-authored content.
|
|
const size = statSync(tarball).size;
|
|
await uploadFileToS3(s3, tarball, projectS3Uri, "application/gzip");
|
|
return {
|
|
siteId,
|
|
bucketName: opts.bucketName,
|
|
projectS3Uri,
|
|
bytes: size,
|
|
uploadedAt: new Date().toISOString(),
|
|
uploaded: true,
|
|
};
|
|
} finally {
|
|
rmSync(workdir, { recursive: true, force: true });
|
|
}
|
|
}
|
|
|
|
async function headObject(
|
|
s3: S3Client,
|
|
bucket: string,
|
|
key: string,
|
|
): Promise<{ bytes: number; lastModified: string } | null> {
|
|
try {
|
|
const res = await s3.send(new HeadObjectCommand({ Bucket: bucket, Key: key }));
|
|
return {
|
|
bytes: typeof res.ContentLength === "number" ? res.ContentLength : 0,
|
|
lastModified:
|
|
res.LastModified instanceof Date
|
|
? res.LastModified.toISOString()
|
|
: new Date().toISOString(),
|
|
};
|
|
} catch (err) {
|
|
// The SDK throws different error shapes for 404 vs 403 vs network;
|
|
// a 404 means "needs upload" and is the most common case. Anything
|
|
// else propagates so callers see auth / network failures.
|
|
const status = (err as { $metadata?: { httpStatusCode?: number } }).$metadata?.httpStatusCode;
|
|
if (status === 404) return null;
|
|
const name = (err as { name?: string }).name;
|
|
if (name === "NotFound" || name === "NoSuchKey") return null;
|
|
throw err;
|
|
}
|
|
}
|