Disaster Recovery
PRISM is stateless by design, which makes its DR story short and honest.
What can be lost
Nothing durable lives in PRISM. The render cache is in-memory and rebuildable from the origin at any time; configuration and licence are files you already manage; there is no database, queue, or writable volume. Losing every replica loses no data — it loses warmth.
Recovery
Redeploy from the container registry or packages and point the edge back. Reference numbers from the nightly CI evidence stack (4-tab pool on shared runners — a deliberately modest environment):
- process start → serving (healthy, proxying): the shipped image healthcheck budgets 60 s; every green e2e and nightly run includes PRISM coming up inside its compose health budget, typically well under it
- first render of a page: the nightly load suite enforces p95 < 3 s on warm fast paths and p95 < 5 s across mixed bot traffic; humans are proxied throughout and never wait on a render
- warmup of a catalogue: roughly
URLs × device classes / pool size × render time— budget minutes, not hours, and run it immediately after recovery
RTO ≈ deploy time + warmup window. RPO = zero — there is no PRISM state to lose. Your origin's DR posture, not PRISM's, is the binding constraint: PRISM down degrades crawlers to your origin's raw responses only if the edge is configured to fail open, so decide that failover explicitly at the CDN.
Reducing the cold window
- Spread replicas (
topologySpreadConstraintsin the Helm chart) so a zone loss keeps warm replicas serving. - Keep a warmup URL list (sitemap) versioned next to your deployment and run it post-rollout — every deploy is a small DR drill.
- The roadmap's shared cache tier will carry warmth across replicas; until then, warmth is per-replica and the warmup is the recovery procedure.