Capacity Planning
How many replicas, how many tabs, how much memory. Every number here comes from a cited source — the nightly evidence suite, the scale matrix, or a constant in the code — and the formulas are written so you can substitute your own measurements, because your page weight dominates every one of them.
The two regimes
PRISM serves two kinds of request with wildly different costs:
- Warm (cache hit): served from memory, no Chrome involved. On the
nightly reference stack the whole client-side request is sub-5ms at p95 —
that is k6's
http_req_duration, so it bounds the hit path rather than measuring it directly, and it includes the client's overhead. A replica's warm throughput is bounded by proxy CPU, and it is high — the reference stack sustains its mixed load without the hit path ever being the constraint. - Cold (render): a Chrome navigation, seconds not milliseconds. This is the resource you are actually planning: render throughput ≈ tabs ÷ average render seconds.
Sizing is therefore about one question: how many renders per second must this fleet sustain?
Memory
replica memory ≈ 400 MiB + 525 MiB × tabs
The constants come from memory_budget_warning in the code, and PRISM
enforces them as a startup warning when the container limit cannot fit the
configured pool. They are calibrated against the nightly scale matrix, which
renders deliberately adversarial pages and settles a 4-tab pool at
2.4–2.6 GiB — not against the gentler pool soak, because a guard sized to
the lighter of two numbers we measure ourselves goes quiet exactly when pages
get heavy.
What the shipped artefacts actually set: the Kubernetes manifest requests 4 Gi and limits at 6 Gi, and its config runs 4 tabs. The code's own default is 8 tabs, which by the formula wants ~4.6 GiB — so a 4 Gi limit with the code default now produces a startup warning rather than an OOM kill under load.
Rules of thumb:
- 4 tabs → ~2.5 GiB (request 4 Gi, limit 6 Gi as shipped); 8 tabs → ~4.6 GiB, so 6–8 Gi.
- Chrome under memory pressure does not degrade, it gets OOM-killed and takes every in-flight render with it — budget for your heaviest pages, not your average ones.
Render throughput
renders/sec per replica ≈ tabs ÷ render_seconds
render_seconds is the number you must measure for your own site — page
weight and origin latency dominate it. Two grounded reference points:
- Uncontended reference environment: first render of a typical SPA page in the low seconds (the nightly load suite enforces p95 < 3s on warm fast paths and < 5s across mixed bot traffic).
- Fully contended: the scheduled nightly's mixed-traffic run — 30 virtual
users against a 4-tab pool sharing one CI runner with the load generator
and the origin. PRISM's own request logs in that run put cold renders at a
median of 9.2s across 20 of them. k6 does not report that number; it comes
from
render_time_msin the log lines, which is why the sources table says "derived". Treat it as the pessimistic bound, and as a small sample.
With render_seconds = 3 and 4 tabs, a replica sustains ~1.3 renders/s;
at 8 tabs, ~2.7/s.
The warmth budget
A cache entry lives cache.ttl seconds, so keeping a corpus warm costs:
required renders/sec = unique URLs × device classes ÷ ttl seconds
(device classes is 2 when viewport-aware rendering is enabled — though a
crawler's request warms both classes at once, so steady-state crawl traffic
does part of this work for free.)
Worked example: 10,000 URLs × 2 classes ÷ 3600s TTL ≈ 5.6 renders/s sustained. At 3s renders and 8 tabs per replica (~2.7 renders/s each), that is 3 replicas to hold warmth with a little headroom — before crawl bursts.
Crawl bursts
Googlebot does not spread its visits politely across your TTL window.
Bursts land on the queue; prism_render_queue_wait_ms (p95 alerted at 5s
by the ops pack) tells you when the pool is undersized for them. The
levers, in order:
- Warmup after every deploy and scale-up — a cold replica converts crawler hits into renders at the worst moment.
- Longer TTLs with grace — serving slightly stale to a crawler beats queueing it.
- More replicas — horizontal first: more tabs per replica costs Chrome memory linearly and CPU contention non-linearly.
- Verified bots with enforcement — the queue defence against traffic that merely claims to be a crawler; see the bot configuration page.
What not to size for
- Cache hit ratio is an outcome, not an input — the nightly mixed- traffic scenario converges to ~93% on a 20-URL corpus in 3 minutes. Size for the render rate; the ratio follows from TTL × corpus × crawl pattern.
- Human traffic in bot-only mode is proxied, not rendered — it costs proxy CPU only and does not enter the render budget.
Sources
| Number | Source |
|---|---|
| 400 + 525×tabs MiB | memory_budget_warning constants in src/config.rs, calibrated to the nightly scale matrix and enforced at startup; a test fails if this table and the code disagree |
| 2.4–2.6 GiB @ 4 tabs | Scale matrix full tier, scheduled nightly 2026-08-24 (adversarial pages) |
| warm p95 < 5ms (whole request), 93% hit ratio | k6 cache_mixed_traffic, scheduled nightly 32687149134 — http_req_duration and cache_hits/cache_total; k6 reports neither a cache-hit-only latency nor a cold-render figure |
| cold p50 9.2s (contended) | Derived, not reported: the render_time_ms field of PRISM's own request logs in that run's container output, 20 cold renders, median. Reproduce with grep -oE 'render_time_ms=[0-9]+' over the job log |
| p95 < 3s / < 5s enforced | k6 thresholds the nightly suite fails on |