Skip to main content

Capacity Planning

How many replicas, how many tabs, how much memory. Every number here comes from a cited source — the nightly evidence suite, the scale matrix, or a constant in the code — and the formulas are written so you can substitute your own measurements, because your page weight dominates every one of them.

The two regimes​

PRISM serves two kinds of request with wildly different costs:

  • Warm (cache hit): served from memory, no Chrome involved. On the nightly reference stack the whole client-side request is sub-5ms at p95 — that is k6's http_req_duration, so it bounds the hit path rather than measuring it directly, and it includes the client's overhead. A replica's warm throughput is bounded by proxy CPU, and it is high — the reference stack sustains its mixed load without the hit path ever being the constraint.
  • Cold (render): a Chrome navigation, seconds not milliseconds. This is the resource you are actually planning: render throughput ≈ tabs ÷ average render seconds.

Sizing is therefore about one question: how many renders per second must this fleet sustain?

Memory​

replica memory ≈ 400 MiB + 525 MiB × tabs

The constants come from memory_budget_warning in the code, and PRISM enforces them as a startup warning when the container limit cannot fit the configured pool. They are calibrated against the nightly scale matrix, which renders deliberately adversarial pages and settles a 4-tab pool at 2.4–2.6 GiB — not against the gentler pool soak, because a guard sized to the lighter of two numbers we measure ourselves goes quiet exactly when pages get heavy.

What the shipped artefacts actually set: the Kubernetes manifest requests 4 Gi and limits at 6 Gi, and its config runs 4 tabs. The code's own default is 8 tabs, which by the formula wants ~4.6 GiB — so a 4 Gi limit with the code default now produces a startup warning rather than an OOM kill under load.

Rules of thumb:

  • 4 tabs → ~2.5 GiB (request 4 Gi, limit 6 Gi as shipped); 8 tabs → ~4.6 GiB, so 6–8 Gi.
  • Chrome under memory pressure does not degrade, it gets OOM-killed and takes every in-flight render with it — budget for your heaviest pages, not your average ones.

Render throughput​

renders/sec per replica ≈ tabs ÷ render_seconds

render_seconds is the number you must measure for your own site — page weight and origin latency dominate it. Two grounded reference points:

  • Uncontended reference environment: first render of a typical SPA page in the low seconds (the nightly load suite enforces p95 < 3s on warm fast paths and < 5s across mixed bot traffic).
  • Fully contended: the scheduled nightly's mixed-traffic run — 30 virtual users against a 4-tab pool sharing one CI runner with the load generator and the origin. PRISM's own request logs in that run put cold renders at a median of 9.2s across 20 of them. k6 does not report that number; it comes from render_time_ms in the log lines, which is why the sources table says "derived". Treat it as the pessimistic bound, and as a small sample.

With render_seconds = 3 and 4 tabs, a replica sustains ~1.3 renders/s; at 8 tabs, ~2.7/s.

The warmth budget​

A cache entry lives cache.ttl seconds, so keeping a corpus warm costs:

required renders/sec = unique URLs × device classes ÷ ttl seconds

(device classes is 2 when viewport-aware rendering is enabled — though a crawler's request warms both classes at once, so steady-state crawl traffic does part of this work for free.)

Worked example: 10,000 URLs × 2 classes ÷ 3600s TTL ≈ 5.6 renders/s sustained. At 3s renders and 8 tabs per replica (~2.7 renders/s each), that is 3 replicas to hold warmth with a little headroom — before crawl bursts.

Crawl bursts​

Googlebot does not spread its visits politely across your TTL window. Bursts land on the queue; prism_render_queue_wait_ms (p95 alerted at 5s by the ops pack) tells you when the pool is undersized for them. The levers, in order:

  1. Warmup after every deploy and scale-up — a cold replica converts crawler hits into renders at the worst moment.
  2. Longer TTLs with grace — serving slightly stale to a crawler beats queueing it.
  3. More replicas — horizontal first: more tabs per replica costs Chrome memory linearly and CPU contention non-linearly.
  4. Verified bots with enforcement — the queue defence against traffic that merely claims to be a crawler; see the bot configuration page.

What not to size for​

  • Cache hit ratio is an outcome, not an input — the nightly mixed- traffic scenario converges to ~93% on a 20-URL corpus in 3 minutes. Size for the render rate; the ratio follows from TTL × corpus × crawl pattern.
  • Human traffic in bot-only mode is proxied, not rendered — it costs proxy CPU only and does not enter the render budget.

Sources​

NumberSource
400 + 525×tabs MiBmemory_budget_warning constants in src/config.rs, calibrated to the nightly scale matrix and enforced at startup; a test fails if this table and the code disagree
2.4–2.6 GiB @ 4 tabsScale matrix full tier, scheduled nightly 2026-08-24 (adversarial pages)
warm p95 < 5ms (whole request), 93% hit ratiok6 cache_mixed_traffic, scheduled nightly 32687149134 — http_req_duration and cache_hits/cache_total; k6 reports neither a cache-hit-only latency nor a cold-render figure
cold p50 9.2s (contended)Derived, not reported: the render_time_ms field of PRISM's own request logs in that run's container output, 20 cold renders, median. Reproduce with grep -oE 'render_time_ms=[0-9]+' over the job log
p95 < 3s / < 5s enforcedk6 thresholds the nightly suite fails on