Skip to main content

Render Configuration

The [render] section controls how PRISM renders pages using a headless Chromium browser. It includes sub-sections for the browser tab pool, circuit breaker, HTML postprocessing, viewport emulation, and content validation.

TOML Example​

[render]
wait_for = "load"
timeout_secs = 10
block_resources = ["font", "image", "media", "stylesheet"]
status_from_meta = false
fail_on_console_errors = false
min_html_size = 512
post_wait_js = "document.querySelector('.lazy')?.classList.remove('hidden')"

[render.pool]
tabs = 8
max_renders_per_tab = 50
queue_max = 100

[render.circuit_breaker]
failure_threshold = 5
recovery_timeout_secs = 30
half_open_max_requests = 1

[render.postprocess]
enabled = true
strip_scripts = true
strip_noscript = true
strip_comments = true
strip_event_handlers = true
strip_hydration_attrs = true
resolve_lazy_images = true

[render.viewport]
enabled = false
mobile_width = 412
mobile_height = 915
desktop_width = 1920
desktop_height = 1080

[render.content_validation]
enabled = true
min_text_length = 100
require_title = true
min_html_bytes = 1024

Top-Level Parameters​

ParameterTypeDefaultDescription
wait_forString"load"Page readiness signal: load, domcontentloaded, networkidle, or selector:<CSS>
timeout_secsInteger10Maximum seconds to wait for a page render. Must be > 0
block_resourcesArray of Strings["font", "image", "media", "stylesheet"]Resource types to block during rendering
status_from_metaBooleanfalseExtract HTTP status from <meta name="render:status_code"> or <meta name="prism:status">
fail_on_console_errorsBooleanfalseFall back to origin if JS console errors detected and HTML is below min_html_size
min_html_sizeInteger512Minimum rendered HTML bytes; below this with console errors triggers fallback
post_wait_jsString or nullnullJavaScript to execute after wait_for completes but before HTML extraction
auto_scrollBooleantrueScroll the page before extracting so intersection observers fire and lazy-loaded content is present. See auto_scroll
accept_invalid_certsBooleanfalseLet Chrome render pages whose TLS certificate does not validate. See accept_invalid_certs

wait_for​

Controls when PRISM considers the page "ready" to extract HTML:

  • load (default): Waits for the browser load event plus a small (~100 ms) settle window. A real readiness signal — fires after the document and all sub-resources finish loading — but it does not wait for client-side fetches that hydrate the SPA after page load. Good for static-ish SPAs where the meaningful content is server-rendered or already in the bundle.
  • domcontentloaded: Fires when the HTML is parsed, before sub-resources finish. Faster than load, almost always too early for hydration-heavy SPAs.
  • networkidle: Waits until there are no more than 2 in-flight network connections for 500 ms. Catches most XHR/fetch hydration but is heuristic — long-polling, analytics beacons, or sticky websockets can defeat it (extending render time to the full timeout_secs).
  • selector:<CSS>: Waits until a CSS selector appears in the DOM. Example: selector:#app-loaded or selector:[data-hydrated="true"]. Recommended for production — the SPA explicitly tells PRISM when it is ready, so render time is bounded by your app's actual hydration, not by a network heuristic.

:::tip Production recommendation For sites where SEO matters, add an app-ready marker to your SPA root component (e.g. <div id="root" data-app-ready="true"> once your data is loaded) and set wait_for = "selector:[data-app-ready]". This eliminates "rendered too early — hero JSON missing" failures that are otherwise hard to diagnose because they only happen under load. :::

status_from_meta​

When enabled, PRISM scans rendered HTML for meta tags like:

<meta name="render:status_code" content="404">
<meta name="prism:status" content="404">

This lets SPAs that render 404 pages client-side communicate the correct HTTP status code to crawlers.

post_wait_js​

Optional JavaScript executed after the wait_for condition is met but before HTML is extracted. Useful for DOM fixups on SPAs with lazy hydration, such as removing hidden classes that a framework has not yet removed.

auto_scroll​

Many SPAs only populate a product grid once it scrolls into view, so without this pass they serialise empty. PRISM scrolls the page in steps before extracting, bounded to roughly five seconds and 200 steps.

The cost is that on a tall page it spends that entire budget on every render. A site that hydrates everything up front gains nothing and should turn it off:

[render]
auto_scroll = false

Leave it on if any part of the site lazy-loads content on scroll — turning it off there produces rendered pages with empty sections, which is worse than a slow render.

accept_invalid_certs​

Off by default: Chrome refuses expired, self-signed and hostname-mismatched certificates, on the origin and on every subresource a page pulls in.

:::danger This weakens a real boundary PRISM renders attacker-influenced content and then caches the result and serves it to crawlers as your site's own. Accepting a bad certificate lets anyone on the network path substitute a page or a subresource and have PRISM publish it under your name, with no TLS warning anywhere because nothing is listening for one. :::

Turn it on only for an internal origin whose certificate you control and have deliberately not had signed — and prefer adding your CA to the trust store instead:

[render]
accept_invalid_certs = true

Pool Configuration​

The [render.pool] sub-section controls the Chromium browser tab pool.

ParameterTypeDefaultDescription
tabsInteger8Number of browser tabs to keep open. Must be > 0
max_renders_per_tabInteger50Deprecated and ignored. Each render already gets a fresh BrowserContext that is disposed afterwards, so per-tab recycling has no work left to do. Accepted so existing configs keep loading. Use max_browser_age_secs / max_renders_per_browser
queue_maxInteger100Maximum queued render requests before rejecting new ones
max_browser_age_secsInteger3600Recycle the Chrome process once it has been up this long. 0 disables
max_renders_per_browserInteger5000Recycle the Chrome process after this many renders. 0 disables

Increasing tabs improves concurrent throughput but uses more memory. Each tab consumes roughly 50-100 MB of RAM depending on page complexity.

Circuit Breaker Configuration​

The [render.circuit_breaker] sub-section implements the circuit breaker pattern to protect against cascading failures when Chromium is unhealthy.

ParameterTypeDefaultDescription
failure_thresholdInteger5Consecutive failures before opening the circuit
recovery_timeout_secsInteger30Seconds to wait before attempting recovery (half-open state)
half_open_max_requestsInteger1Number of test requests allowed in half-open state

When the circuit is open, all render requests immediately fall back to the origin response instead of waiting for Chromium.

Postprocess Configuration​

The [render.postprocess] sub-section controls HTML cleanup applied to rendered output before caching and serving.

ParameterTypeDefaultDescription
enabledBooleantrueEnable or disable all postprocessing
strip_scriptsBooleantrueRemove <script> tags from rendered HTML
strip_noscriptBooleantrueRemove <noscript> tags
strip_commentsBooleantrueRemove HTML comments
strip_event_handlersBooleantrueRemove inline event handlers (onclick, onload, etc.)
strip_hydration_attrsBooleantrueRemove framework hydration attributes (data-reactid, ng-*, etc.)
resolve_lazy_imagesBooleantrueMove data-src to src for lazy-loaded images

Postprocessing produces cleaner HTML for search engine crawlers and reduces response size.

Viewport Configuration​

The [render.viewport] sub-section enables viewport emulation for mobile-aware rendering.

ParameterTypeDefaultDescription
enabledBooleanfalseEnable viewport-aware rendering
mobile_widthInteger412Width in pixels for mobile viewport
mobile_heightInteger915Height in pixels for mobile viewport
desktop_widthInteger1920Width in pixels for desktop viewport
desktop_heightInteger1080Height in pixels for desktop viewport

When enabled, PRISM detects the requesting device type and sets the browser viewport accordingly. This ensures responsive SPAs render the correct layout for mobile vs. desktop crawlers (e.g., Googlebot Mobile).

Content Validation Configuration​

The [render.content_validation] sub-section validates rendered output quality before caching and serving.

ParameterTypeDefaultDescription
enabledBooleantrueEnable content validation
min_text_lengthInteger100Minimum visible text characters (after stripping HTML tags)
require_titleBooleantrueRequire a non-empty <title> tag in rendered output
min_html_bytesInteger1024Minimum total HTML size in bytes

Content validation catches broken renders (empty pages, error screens) and prevents them from being cached or served to crawlers.

Example Use Cases​

React SPA with slow hydration​

[render]
wait_for = "selector:#app-ready"
timeout_secs = 15
post_wait_js = "document.querySelectorAll('.skeleton').forEach(el => el.remove())"

[render.pool]
tabs = 12
queue_max = 200

High-traffic site with conservative settings​

[render]
wait_for = "load"
timeout_secs = 8
fail_on_console_errors = true
min_html_size = 1024

[render.circuit_breaker]
failure_threshold = 3
recovery_timeout_secs = 60

[render.content_validation]
enabled = true
min_text_length = 200
min_html_bytes = 2048

Minimal postprocessing for debugging​

[render.postprocess]
enabled = true
strip_scripts = false
strip_comments = false
strip_hydration_attrs = false