Skip to main content

Disaster Recovery

PRISM is stateless by design, which makes its DR story short and honest.

What can be lost​

Nothing durable lives in PRISM. The render cache is in-memory and rebuildable from the origin at any time; configuration and licence are files you already manage; there is no database, queue, or writable volume. Losing every replica loses no data — it loses warmth.

Recovery​

Redeploy from the container registry or packages and point the edge back. Reference numbers from the nightly CI evidence stack (4-tab pool on shared runners — a deliberately modest environment):

  • process start → serving (healthy, proxying): the shipped image healthcheck budgets 60 s; every green e2e and nightly run includes PRISM coming up inside its compose health budget, typically well under it
  • first render of a page: the nightly load suite enforces p95 < 3 s on warm fast paths and p95 < 5 s across mixed bot traffic; humans are proxied throughout and never wait on a render
  • warmup of a catalogue: roughly URLs × device classes / pool size × render time — budget minutes, not hours, and run it immediately after recovery

RTO ≈ deploy time + warmup window. RPO = zero — there is no PRISM state to lose. Your origin's DR posture, not PRISM's, is the binding constraint: PRISM down degrades crawlers to your origin's raw responses only if the edge is configured to fail open, so decide that failover explicitly at the CDN.

Reducing the cold window​

  • Spread replicas (topologySpreadConstraints in the Helm chart) so a zone loss keeps warm replicas serving.
  • Keep a warmup URL list (sitemap) versioned next to your deployment and run it post-rollout — every deploy is a small DR drill.
  • The roadmap's shared cache tier will carry warmth across replicas; until then, warmth is per-replica and the warmup is the recovery procedure.