Headless Chromium renders caller-supplied HTML on two paths, and neither applied the Content-Security-Policy the app packager injects for exports. Inline script in a preview scene or an uploaded render project could reach loopback and internal addresses, and on the render path the response is painted into the returned MP4. - Add a single untrusted-HTML CSP and an injector that places the policy as the first node the parser processes (ASCII whitespace only; a leading BOM is stripped; a byte-level injector preserves non-UTF-8 bodies). - Inject the policy into the interactive preview srcDoc and add a Puppeteer request guard that blocks non-data/blob/about requests from the untrusted frame and non-about main-frame navigations. - Harden every extracted project HTML, and sanitize framed same-origin .svg and .xhtml documents (scripts, on* handlers, javascript: URLs, foreignObject, nested frames removed via parse5), which cannot carry a meta CSP. - Document the residual top-level-navigation risk on the render path, which the egress lockdown must contain. - Add real-Chromium boundary tests (zero listener hits incl. WebSocket) and a packager-equivalence test for the policy. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
28 KiB
@openmaic/render-service
Isolated MP4 render service for OpenMAIC's classroom video export (issue #866).
The main app compiles a classroom to a self-contained Hyperframes project ZIP
(index.html + assets/ + vendored GSAP) entirely in the browser. This service
takes that ZIP and renders it to an MP4 with @hyperframes/producer, which
drives headless Chromium (frame capture) + FFmpeg (encode). It runs in its own
Node 22 container because the producer needs Node ≥ 22, Chromium, and FFmpeg —
none of which belong in the Next.js runtime.
It is an opt-in capability: when the app has no RENDER_SERVICE_URL
configured, in-app export degrades to downloading the project ZIP for local CLI
rendering. Nothing here is required for the app to run.
HTTP API
Rendering is asynchronous (a 10-minute video can take tens of minutes): submit, poll, then download. Job ids are opaque.
| Method + path | Purpose |
|---|---|
POST /render |
multipart: project (the ZIP) + fps, quality, format fields → 202 { jobId } |
POST /preview |
JSON: persisted scene, stage context, and viewport → synchronous PNG |
GET /render/:jobId |
status/progress plus actual capture, worker, profile, and runtime metrics |
GET /render/:jobId/download |
stream the MP4 (or 302 to a presigned URL) once succeeded |
DELETE /render/:jobId |
cancel a queued/running job |
GET /health |
resource profile, observed runtime versions, and whether renders are being admitted |
status is one of queued | running | succeeded | failed | cancelled;
progress is 0..1.
Admission: 429s and /health
Admission rejections are machine-readable. POST /render answers 429 with a
JSON body { error, reason } where reason is one of:
queue_full— the deployment-wide capRENDER_MAX_QUEUE(reserved + queued + running jobs) is exhausted; back off and retry.per_identity_limit— this client identity already holdsRENDER_MAX_JOBS_PER_USERactive renders.
POST /preview also answers 429 with { error, reason }, where reason is
one of:
preview_queue_full—RENDER_PREVIEW_MAX_IN_FLIGHTpreviews are already live (buffering or executing).preview_per_user_limit— this client identity already holdsRENDER_PREVIEW_MAX_PER_USERpreviews.capacity_busy— the shared execution slot is occupied, including by a video render. Previews never queue for this slot; retry later.
GET /health reports accepting: boolean — whether the queue cap currently has
room for another video render — it reflects the render queue only, not
preview admission or execution-slot availability. Callers and load balancers
can shed traffic before probing with a real render. The flag is deliberately
aggregate-only: no queue depths or occupancy counts, and never any per-identity
data. Identity keys are client IPs when the service runs behind
TRUST_PROXY_HEADERS=true, so exposing the per-identity map would publish the
IP address of everyone currently rendering.
Previews share the service's single Chromium execution slot with video renders,
but never wait in the video queue. Once a preview has passed its body and
semantic checks, it renders immediately or fails with capacity_busy.
/preview requires fully self-contained scenes (inline code and data: URLs
only); network, blob:, and relative references are rejected. Slide media must
use data: URLs, and interactive scenes must contain embedded HTML whose
resource references are inline or data: URLs. Non-self-contained inputs fail
with 422 instead of returning a misleading PNG. Detection is best-effort:
some exotic external constructs (for example, object data, track src, and
CSS @import) may pass validation; self-contained scenes are the caller's
responsibility, and end-to-end fidelity for persisted scenes arrives with
caller-side preparation (tracked separately).
Environment
| Var | Default | Meaning |
|---|---|---|
PORT |
9000 |
Listen port. |
RENDER_RESOURCE_PROFILE |
standard |
standard prefers BeginFrame, permits producer compatibility fallback to screenshot, and requires 8 GiB; low-memory forces screenshot capture and requires 4 GiB. Both fix one producer worker, one render, and one extraction. |
RENDER_MAX_CONCURRENCY |
profile: 1 |
Must match the selected profile. Renders beyond the single execution slot queue FIFO. |
RENDER_MAX_CONCURRENT_EXTRACTIONS |
profile: 1 |
Must match the selected profile; bounds archive expansion to one 512 MiB expanded archive at a time. |
RENDER_MAX_JOBS_PER_USER |
1 |
Active jobs allowed per client identity (0 disables the guard — see note below). |
RENDER_MAX_QUEUE |
20 |
Max jobs in the system (reserved+queued+running) before new submits get 429. |
RENDER_JOB_TTL_MS |
1800000 |
How long finished jobs + artifacts live before cleanup. |
RENDER_JOB_DEADLINE_MS |
2700000 |
Hard per-job wall-clock deadline; overruns are aborted and marked failed. |
RENDER_PREVIEW_TIMEOUT_MS |
20000 |
Hard wall-clock deadline for a synchronous preview, including body parsing and Chromium cleanup. |
RENDER_PREVIEW_MAX_IN_FLIGHT |
8 |
Maximum admitted previews across buffering and execution. |
RENDER_PREVIEW_MAX_PER_USER |
2 |
Concurrent previews per owner identity; 0 disables the guard for deployments whose preview callers do not supply an owner identity (see note below). |
RENDER_PREVIEW_MAX_JSON_BYTES |
33554432 |
Maximum preview JSON body size (32 MiB), enforced on declared length and streamed bytes independently of the ZIP upload cap. |
RENDER_CHUNK_EXECUTION |
false |
Opt in to the bounded local plan -> renderChunk -> assemble executor. The HTTP API stays unchanged. |
RENDER_CHUNK_COUNT |
1 |
Number of deterministic closed-GOP chunks planned for a render when chunk execution is enabled. |
RENDER_CHUNK_WORKERS |
profile: 1 |
Maximum producer capture workers inside one chunk. This remains explicit so chunk fan-out cannot multiply nested producer workers. |
RENDER_MAX_PARALLEL_CHUNKS |
1 |
Maximum chunks executed concurrently in the local process. Chunks beyond the bound wait locally. |
RENDER_CHUNK_SIZE_FRAMES |
unset | Optional fixed frame count per planned chunk. |
RENDER_TARGET_CHUNK_FRAMES |
unset | Optional target frame count used by the producer planner when deriving chunk boundaries. |
RENDER_MAX_UPLOAD_BYTES |
314572800 |
Max compressed archive size accepted (300 MB); enforced on real bytes, before buffering. |
RENDER_MAX_ENTRIES |
5000 |
Max entries allowed in the archive. |
RENDER_MAX_ENTRY_BYTES |
209715200 |
Max expanded size of any single entry (200 MB). |
RENDER_MAX_EXPANDED_BYTES |
536870912 |
Max total expanded size across all entries (512 MB). |
RENDER_MAX_COMPRESSION_RATIO |
200 |
Max expanded:compressed ratio per entry (ZIP-bomb guard). |
RENDER_EGRESS_LOCKDOWN |
true |
Install the iptables egress lockdown at startup (needs root + CAP_NET_ADMIN); fails closed — the container exits if the rules can't be applied. Set false to run unisolated. |
PRODUCER_TMP_PROJECT_DIR |
/tmp/openmaic-renders |
Scratch dir for unzipped projects + outputs. |
PRODUCER_BROWSER_GPU_MODE |
profile-controlled | Both profiles use the software/SwiftShader selector; standard keeps BeginFrame eligible with PRODUCER_FORCE_SCREENSHOT=false, while low-memory forces screenshot capture. This is not a host GPU requirement. Do not override it directly. |
PRODUCER_LOW_MEMORY_MODE |
profile-controlled | Explicitly false for standard and true for low-memory; cgroup heuristics cannot silently switch the selected profile. |
PRODUCER_MAX_WORKERS |
profile: 1 |
Explicit for both supported profiles so producer auto-sizing cannot raise the worker count. |
PRODUCER_ENABLE_BROWSER_POOL |
profile: false |
Disabled because both supported profiles use one worker; no additional Chromium instances are admitted. |
PRODUCER_HEADLESS_SHELL_PATH |
/usr/bin/chromium-headless-shell (container) |
Chromium headless shell executable used by producer's beginFrame resolver. Regular Chromium is not equivalent: it may resolve as beginFrame-capable and then reject HeadlessExperimental.beginFrame, causing a screenshot fallback. |
RENDER_REQUIRE_BEGINFRAME |
profile-controlled | false for both supported profiles. standard requests BeginFrame but accepts producer compatibility fallback; low-memory forces screenshot. This internal compatibility knob must not be overridden independently. |
PRODUCER_PUPPETEER_PROTOCOL_TIMEOUT_MS |
900000 (set in Compose) |
CDP timeout headroom for long frame ranges. The producer default of 300 seconds caused long jobs to fall back from four workers to two. |
HF_STATIC_DEDUP |
false (set in Compose) |
Temporary OpenMAIC-export workaround: these long slide compositions currently exhaust producer's 15-second verification budget and disable dedup anyway. Skipping the doomed verification removes the fixed startup cost without changing frames. |
RENDER_HOME |
/app |
Writable home used after the entrypoint drops privileges. Producer font caches live under $RENDER_HOME/.cache, never /root/.cache. |
PUPPETEER_EXECUTABLE_PATH |
/usr/bin/chromium-headless-shell |
System Chromium headless shell (set in the image). |
Chunk settings are profile-bounded: the standard profile allows at most one producer worker per chunk and four concurrent chunks; low-memory allows one of each. Values above the selected profile limit fail startup rather than silently multiplying browser, FFmpeg, and temporary-disk usage.
Client identity for the per-user guard is taken from the x-openmaic-client
header, which the app's proxy sets. A client-supplied userId form field is
ignored. The app derives that header from x-forwarded-for/x-real-ip only
when the operator sets TRUST_PROXY_HEADERS=true (and a real reverse proxy
overwrites those headers); otherwise all callers share one direct identity, so
the default directly-exposed Compose topology can't be gamed by spoofing
forwarding headers.
Per-user guard vs. shared identity. Without
TRUST_PROXY_HEADERS=trueevery caller collapses into one shareddirectidentity bucket, so any nonzeroRENDER_MAX_JOBS_PER_USERthrottles the whole deployment to that many concurrent renders — not just one user. This is why the defaultdocker-compose.ymlsetsRENDER_MAX_JOBS_PER_USER=0(guard off) and relies onRENDER_MAX_CONCURRENCY+RENDER_MAX_QUEUE(see the comment beside that variable in the Compose file). Enable the per-user guard only behind a trusted proxy that supplies a real per-user identity.Preview callers are different: the app sends the durable session owner id in
x-openmaic-client, soRENDER_PREVIEW_MAX_PER_USERstays enabled and defaults to 2. Set it to 0 only for deployments that call/previewwithout an owner identity.
Security / isolation
The uploaded archive is untrusted, so extraction is bounded before any bytes are decompressed (entry count, per-entry and total expanded size, and compression ratio — see the limits above), guarding against ZIP bombs. Extraction runs on fflate's worker (off the event loop) and is concurrency-capped so admitted jobs can't stack the per-archive RAM ceiling.
The composition HTML is then executed in headless Chromium. Two boundaries keep that untrusted page contained:
- No inbound-to-app bridge. The container's entrypoint installs an iptables
egress lockdown (drop all outbound except loopback + replies on app-initiated
connections), so Chromium can't open connections back to the app — even though
they share the Compose network so the app can reach the service. This needs the
container to run with
CAP_NET_ADMIN(cap_add: [NET_ADMIN], already set in the Compose file). WithRENDER_EGRESS_LOCKDOWN=true(the default) the entrypoint fails closed: if the rules can't be applied (missing capability, backend mismatch) the container exits non-zero rather than start an unisolated service the app would still advertise as healthy. An operator who knowingly accepts an unisolated standalone setup opts out withRENDER_EGRESS_LOCKDOWN=false.scripts/egress-smoke.sh <image>asserts the boundary end-to-end (lockdown active, loopback works, a new outbound connection is blocked). - No internet. In Compose the
rendernetwork isinternal: true(no host or internet gateway). The export ZIP bundles every asset (and GSAP) at build time, so the render needs no outbound at all.
When running standalone, place the service on an isolated network yourself (and keep the egress lockdown on, or accept the risk with the toggle) — it needs no outbound access.
Residual risk
The injected CSP plus request interception close script, fetch/XHR, WebSocket, image, frame and form egress from the untrusted documents. Two channels remain:
- Top-level navigation on
/render.location, a<meta http-equiv=refresh>,window.openandtarget=_topcan replace the top frame; CSP has nonavigate-todirective, and the producer's browser exposes no request interception we can install. The service therefore relies on the container's egress lockdown: withRENDER_EGRESS_LOCKDOWN=false, a rendered page can navigate to any reachable address and the result appears in the output video. Do not run standalone with the lockdown disabled. - Declarative subresources inside a framed SVG/XHTML. Sanitizing removes
script and event handlers, but a surviving
<image href>(or CSSurl()) can still load a subresource through a framed document that has no CSP of its own. The lockdown is what blocks that egress in supported deployments.
Run
Docker (recommended)
The root docker-compose.yml wires this service under the video-export
profile and points the app at it:
docker compose --profile video-export up --build
Standalone (development)
Requires Node 22, Chromium's old headless shell, and FFmpeg on PATH. The
standard profile checks for 8 GiB of available host/cgroup memory before
listening:
cd render-service
npm install
PUPPETEER_EXECUTABLE_PATH=$(which chromium-headless-shell) \
PRODUCER_HEADLESS_SHELL_PATH=$(which chromium-headless-shell) \
RENDER_RESOURCE_PROFILE=standard \
PRODUCER_PUPPETEER_PROTOCOL_TIMEOUT_MS=900000 \
HF_STATIC_DEDUP=false \
npm start
Resource profiles
The default standard CPU profile is the intended 1080p / 30 fps / standard
quality path: it prefers BeginFrame, while allowing producer to select screenshot
for compatibility-sensitive compositions such as iframe GenUI. Producer workers
are fixed at one, and the service admits one render plus one archive extraction
at a time. It requires at least 8 GiB of host/cgroup memory and an existing
headless shell so ordinary compositions remain BeginFrame-eligible. No host GPU
is required or requested; Chromium uses its software/SwiftShader selector.
Use the safe low-memory profile only when BeginFrame latency is less important
than a smaller memory ceiling. It fixes screenshot capture, one worker, one
render, and one extraction, and requires at least 4 GiB:
RENDER_RESOURCE_PROFILE=low-memory \
RENDER_SERVICE_MEMORY_LIMIT=4g \
docker compose --profile video-export up --build
Both /health and GET /render/:jobId make the selection observable. Health
reports the capture policy, requested mode, worker/concurrency bounds, minimum
memory, and observed Node, producer, Chromium, and FFmpeg versions. A completed
job reports requested versus actual capture mode and worker count with the same
version record, so a standard-profile screenshot fallback remains explicit.
The previous fixed 720p short-sample comparison that motivated these profiles was:
| Capture path | Workers | Result |
|---|---|---|
| screenshot | 1 | 28.7 s baseline |
| BeginFrame | 1 | 16.3 s, about 43% faster |
| BeginFrame | 2 | only a small improvement beyond one worker |
| BeginFrame | 4 | no improvement on four vCPU; higher resource pressure |
These measurements justify the one-worker standard, not a general latency SLA. The full 1080p sample took 696.9 s with one worker, with capture about 75.5% and encode about 20.6% of runtime. Four-worker 1080p/4K and multi-job profiles remain unsupported until their combined correctness and memory bounds are validated.
Output parity must be verified separately for each validated profile before claiming parity acceptance.
Scalability
The service is built with three swap points so it can move from a single OSS host to a horizontally-scaled demo deployment without changing the HTTP contract or the app:
RenderExecutor(src/render-executor.ts) —InProcessExecutoradapts the current HyperFrames producer to stable progress, cancellation, deadline, failure, and performance types. A bounded local or remote executor can replace it without changingRenderCoordinatoror the routes.JobStore(src/job-store.ts) — Part A shipsInMemoryJobStore. ARedisJobStoreimplementing the same interface lets any replica serve poll / download requests.ArtifactStore(src/artifact-store.ts) — Part A shipsLocalDiskArtifactStore(streams through the app proxy). AnS3ArtifactStorewhoselocatereturns a presigned URL makes the download route302the browser straight to object storage, bypassing the proxy.
RenderCoordinator owns admission, queueing, job state, artifact registration,
and cleanup while depending only on those three interfaces.
The opt-in local chunk executor uses @hyperframes/producer/distributed without
adding a queue or changing the HTTP contract. It freezes the producer plan and
runtime versions, bounds chunk fan-out and nested producer workers, verifies
chunk bytes and sidecar hashes before ordered assembly, and reuses only a valid
result for an idempotent retry. The default remains the in-process executor.