Changelog
Every release, and for the ones that change behaviour, what to do about it. Look for the Upgrading section of the version you are moving to — that is the migration guide, and it names the things that are visible from outside PRISM.
1.4.0 — 2026-09-03
The version line was renumbered here. Everything below shipped previously as 1.4.0 through 1.7.1 across eight tags in ten days — a pace that said more about how the work was sequenced than about how much of it there was. Those numbers are withdrawn from the package repositories and the registry; this is the release that replaces them, and it contains all of their substance.
Nothing is lost: each entry is the one written when the change was made. They are grouped by kind rather than by the release they happened to land in.
The git tags for the withdrawn versions still exist and cannot be removed — the tag ruleset refuses deletion. They are source history, not a distribution channel; installs come from apt and the registry.
Upgrading
No configuration change is required. Three corrections can make previously hidden origin behaviour visible:
- Repeated response fields are preserved as repeated fields. An origin sending
two
Cache-Control,X-Robots-Tag,Link,Varyor security-header fields now has both honoured instead of a newline-joined value being misparsed or dropped. - Parameters listed in
cache.strip_query_paramsno longer reach Chrome. They were already absent from the shared key, so allowing them to shape the stored document let the first tracking URL decide what every later crawler saw. - Sitemap indexes are now required to keep child sitemaps on the same scheme, host and effective port, and malformed XML fails the warmup instead of reporting a successful partial run. Move a child sitemap outside that authority onto the index authority before upgrading.
Client-supplied forwarding aliases and rewrite hints such as Forwarded,
CF-Connecting-IP, True-Client-IP, X-Original-URL and X-Rewrite-URL are
now removed before the origin request. PRISM continues to set its own trusted
X-Forwarded-For, X-Real-IP, X-Forwarded-Proto and X-Forwarded-Host.
Nothing needs reconfiguring. One thing changes what a cache in front of PRISM
does, and it is the point of the release: a response PRISM declines to store —
a 401, 403, 407, 429, or a temporary redirect the origin did not grant a
lifetime — is now sent with no-store instead of public, max-age=3600. If
something downstream was absorbing those, it will stop, and requests that were
being answered at the edge will reach PRISM. That is the correct behaviour: a
403 pinned at a CDN is a page Googlebot cannot reach until the TTL expires.
Two further effects are improvements you may notice in metrics rather than
behaviour. Cache hit rates on sites with tracking parameters will rise,
because a URL whose query is entirely stripped now shares the bare URL's key
instead of missing it. And the Age reported on a cache miss now includes the
render, so it will read a few seconds higher than before — the same number a
hit for the same document reports.
Nothing needs reconfiguring. Several things behave differently and all of them are visible from outside PRISM, so they are listed with what to expect rather than left to be discovered.
Redirects that call themselves temporary are no longer cached. A 302, 303
or 307 was kept for the route's TTL; now it is stored only if the origin states
a lifetime for it. On a site that redirects a lot this means more renders and a
lower hit rate — and correct behaviour, because a temporary redirect pinned for
an hour outlives the thing it was temporary for. If yours are stable enough to
cache, have the origin say so with Cache-Control: max-age=....
A write drops what was rendered for the same URL. A successful POST, PUT, PATCH or DELETE now invalidates the cached GET of the URL it addressed, on the replica that handled it. Expect a render on the next crawl of that page. This is local — a fleet still needs the durable purge journal, which is 1.8 work.
Cached entries start at their real age. The clock used to start when the
entry was stored, which was after the render. It now accounts for the origin's
Date, its Age, and the time the fetch and render took. For an origin that
sends Date and no Age, this costs a few seconds off a lifetime measured in
hours — not something you will see. For an origin already behind a CDN, the
Age it reports is now respected rather than partly discarded, which is the
point: PRISM stops handing a response a full lifetime it was never granted.
Errors and refusals are announced honestly. A cached 5xx is advertised for
cache.error_ttl_secs rather than the route's TTL, and PRISM's own 429/5xx
carry no-store. If a CDN in front was absorbing origin outages on the longer
number, raise error_ttl_secs — it now governs both ends.
A new response header. Cache-Status (RFC 9211) is sent alongside
X-Prism-Cache, which is unchanged. Nothing reads it unless you point
something at it.
Two new settings, both defaulting to current behaviour:
cache.honor_no_transform (false) and server.max_request_headers (100,
the value already in force). Neither changes anything unless you set it.
Downstream caches will hold less than they did. Nothing needs changing to adopt this, but three things behave differently and all three are visible from outside PRISM:
- A cached 5xx is announced with the lifetime it is actually stored under
(
error_ttl_secs, default 30s) rather than the route's. If you relied on a CDN in front to absorb origin outages for longer, raisecache.error_ttl_secs— that setting now governs both ends. - 401, 403, 407 and 429 from the origin are no longer stored unless the origin
states a lifetime for them. If your origin returns one of these for pages
you want cached, it must say so with
Cache-Control. - PRISM's own 429/502/503/504 carry
no-store. If a CDN was absorbing PRISM's rate limiting for you, that shielding now happens at PRISM's own limiter rather than downstream — checksecurity.rate_limit_per_ipis where you want it.
No configuration keys were added, removed or changed, and no default moved.
One theme: a cache must not outlive the condition it cached. Three responses that describe a moment — an origin failing, an origin refusing, PRISM itself refusing — were each being kept, or advertised, for the lifetime of a page.
If you installed 1.5.1 from a package, the systemd unit it dropped caps
memory below what the default pool needs. dpkg, rpm and apk treat
/usr/lib/systemd/system/prism.service as a configuration-adjacent file, so
check it after upgrading and make sure it reads MemoryMax=6G. Then
systemctl daemon-reload && systemctl restart prism. Nothing else in this
release needs action.
-
New:
[bot.verify]— forward-confirmed reverse DNS for claimed crawlers. Off by default; every key defaults, so existing configs are unaffected. Turnenabledon to observe andenforceon to act. -
One tightening can refuse startup. An
admin.bearer_tokenthat still reads as a placeholder (CHANGE-ME…) is rejected at load. Onlydeploy/k8s/config.tomlever shipped one; if you copied it, replace the value (openssl rand -hex 32). -
The memory guard is stricter. Chrome is now costed at 525 MiB per tab rather than 375, calibrated to what the nightly scale matrix actually observes. A container limit that used to pass silently may now log a startup warning — that warning is the point, and the fix is more memory or fewer tabs, not a smaller estimate.
-
Cache keys escape
|and%in their URL portion. Keys are in-memory only, so there is nothing to migrate; purge patterns are escaped the same way and keep matching. Patterns written against URLs containing a literal|need no change. -
Kubernetes users: both shipped paths now require a licence Secret (
prism-license) and the chart ships a NetworkPolicy that is on by default. SetnetworkPolicy.originor your renders cannot reach the origin — the chart says so on install. -
Every new config key defaults sanely:
license.recheck_minutes = 60,license.grace_hours = 72,admin.metrics_bearer_tokenunset. No schema breaks. -
One behaviour tightening can refuse startup: a
[cache] varylisting a credential header (cookie,authorization,host,x-prism-*,x-forwarded-*) is now rejected at load. No correct deployment had one — such a value copied session tokens into cache keys and handed them to the renderer — but if yours did, remove the entry before upgrading. -
GitHub Releases are gone as of this version; packages and Docker images are unchanged channels, and each release's SBOM + checksums now live at
pkg.trident-cache.com/prism/provenance/<tag>/.
Nothing requires configuration changes. Three behaviours become visible:
- Purges start working in viewport mode. Single-URL and pattern purges
now remove the
::mobile/::desktopvariants they always claimed to remove. If your workflow relied on purges being no-ops (it did not), purge volume is unchanged — but entries actually disappear now. - Warmup renders each URL once per device class when
render.viewportis enabled — twice the warmup renders, so budget warmup windows accordingly. Admin/renderreports the variants in avariantsarray; its top-level fields are unchanged for single-variant deployments. - Coalesced waiters wait out their leader instead of 503ing. Under load
spikes, requests that previously failed at
timeout+5snow wait as long as the leader may legitimately work. Tail latency replaces spurious 503s.
Six things change behaviour an operator has to plan for. Everything else in this release is a fix that needs nothing.
-
Hit rates fall where the origin said not to cache. Pages carrying
no-cache,max-age=0,s-maxage=0orVary: *are no longer stored at all, and a shorters-maxageormax-agenow caps PRISM's own TTL instead of being overruled by it. Magento sendsno-cacheon more routes than most operators expect.Check what your origin actually sends on the routes you care about before reading the new numbers as a regression. An origin that means a page to be cached should say so, and
s-maxageis the directive aimed at PRISM specifically. -
Size
security.rate_limit_per_ipbefore upgrading. Concurrent requests for an uncacheable URL each render now, where one render used to answer all of them. The default budget is three renders per IP per second, and behind a CDN every visitor resolves to the same edge address unlessserver.trusted_proxieslists your proxy — so one burst on ano-cachepage can exhaust it for the whole site. Settrusted_proxies, raise the budget, or both. -
The admin API now authenticates. Anything calling
/purge/*,/renderor/warmupwithout a credential stops working, including scripts that relied on loopback being open. Setadmin.bearer_tokento a value you control; without one PRISM generates a token per run and logs it at startup, which is secure but changes on every restart.Probes are unaffected —
/health,/readyand/startupnever required a credential, so orchestrators and the Docker healthcheck need no change. -
admin.peersnow requires a configuredadmin.bearer_token. Startup refuses the combination, because a generated token cannot be shared between replicas and every forwarded purge would be rejected. Set the same token on every replica. -
Subresources on hosts PRISM cannot resolve are now blocked. The guard that stops a rendered page reaching private addresses used to wave through any hostname it failed to resolve. A third-party asset whose DNS is slow or briefly down is now refused rather than left to fail on its own.
The cost is bounded on purpose — an unresolved verdict is never cached, so a flaky lookup refuses the one subresource in front of it and nothing more. Only applies to cross-origin subresources, and only while
security.block_private_cidrsis on; your own origin is exempt either way. -
X-Forwarded-Hostis pinned whenserver.public_base_urlis set. A clientHostthat does not match the configured one is replaced with it. If your origin legitimately serves several hostnames through one PRISM instance, leavepublic_base_urlunset — behaviour is then unchanged.
Added
-
GET /explainanswers why PRISM did what it did with a URL. The question it exists for is the one the roadmap says has cost days of back-and-forth on live deployments: why will my page not cache? Until 1.6.0 the honest answer was a header value and a shrug, because the decision was an if-else chain producing string literals and nothing could report on it. The render decision and the cache decision now carry typed reasons, so this is mostly a matter of asking them.Three answers, asked separately. Routing: would this path be rendered, and if not, which condition stopped it — pass
user_agentto ask as a specific crawler. Cache: what is stored right now, how old, how much longer it counts as fresh, whether grace may serve it, and every variant under the same canonical URL. Policy: a hypothesis — given astatusandcache_controlas your origin would send them, would the result be stored, for how long, what would be announced downstream, and if not stored, why not, in a sentence.It renders nothing, fetches nothing, and does not touch the cache it reports on.
RenderCache::inspectreads without promoting the entry in the LRU or counting a hit — a diagnostic that moves the numbers it reports would raise the hit rate of anyone debugging a low one, and there is a test that fails if it ever starts to. -
The release record is published where the documentation sends readers.
versioning.mdsays the changelog is the release record and that each release's Upgrading section is the migration guide. The repository is internal, so that sent a reader to a file they could not open. The changelog is now a page on the documentation site, generated fromCHANGELOG.mdat build time rather than copied — a second hand-maintained changelog diverges within a release or two, and the divergence stays invisible until someone acts on the stale half. -
cache.honor_no_transform, defaulting tofalse— and the default is the decision worth explaining. RFC 9111 §5.2.2.6 tells an intermediary not to transform a representation, which is exactly what replacing a shell with rendered HTML does. PRISM has always ignored it; that is now a stated choice with a switch and a counter rather than an omission.The directive exists to stop intermediaries the origin does not control — a carrier proxy recompressing images, a transcoding gateway. PRISM is not one: the origin's own operator installs it, points it at their own site, and pays for that transformation. Magento and other stacks also emit
no-transforminside boilerplateCache-Controlwith no intention about rendering.The failure modes settle it. Obeying by default means an upgrade silently stops rendering for someone who bought PRISM to render, surfacing weeks later as "our SEO stopped working" with no obvious cause. Not obeying breaks the letter of a directive in a way nobody observes. The second is the cheaper mistake — and
prism_no_transform_ignored_totalcounts every response it applies to, whatever the setting, so an operator who disagrees can see it and flip one key.With it on, the response is served from the origin unchanged and marked
x-prism-fallback: no-transform.no-transformsays nothing about storing, and PRISM does not read it as if it did: an origin that sends it with a lifetime still gets its response cached. -
server.max_request_headers, defaulting to 100 — which is what the HTTP layer would have used anyway. The value is not the point; depending on it without saying so was. PRISM's inbound limits were all inherited defaults, and an upgrade that changed one would have moved this deployment's posture with nothing in the diff to notice.The header bytes stay non-configurable and bounded at 8 KiB plus 4 KiB per header, answering
431. That same ceiling bounds the request target, which is why there is no URL-length setting: the URI is parsed out of the same buffer, so there does not need to be one.tests/inbound_limits.rsboots the real connection builder on a loopback port and checks what a socket refuses — 300 headers under a ceiling of 100, 60 under a configured 20, a two-megabyte request target — rather than asserting what a constant says. It needs no Chrome, no Docker and no origin, so it runs on every commit. -
Cache-Status, the field RFC 9211 defines, sent alongsideX-Prism-Cacherather than instead of it — that header predates the standard and deployments read it. The point is that a CDN, a browser devtools pane or an engineer reading someone else's stack can now see what PRISM did without knowing PRISM exists:prism; hit,prism; fwd=miss,prism; fwd=bypass, and a stale hit asprism; hit; detail=stale— a qualified hit rather thanfwd=stale, because the reader was answered from the cache, which is exactly what serving stale is for.No
ttlparameter: the RFC makes every parameter optional, and the remaining freshness is derivable from theAgeandCache-Controlsent beside it — both of which had to be corrected in this same series before that was true. A third statement of the same number is one more thing that can disagree with the other two. -
A write now drops what PRISM rendered for the same URL. RFC 9111 §4.4 requires a cache to invalidate the target URI when an unsafe method comes back with a non-error status. PRISM did not: a POST that changed a page passed through to the origin, and the rendered copy kept being served to crawlers until its TTL ran out, with nothing in the system aware the page had changed. Every variant is dropped — an exact-key delete would leave the mobile and
Varycopies serving the old page while reporting success — and the URL is resolved by asking the render pipeline, exactly as the admin purge does, so the three ways an entry can be dropped cannot drift apart.Safety is read as RFC 9110 §9.2.1 defines it, a closed set of GET, HEAD, OPTIONS and TRACE, so an extension method counts as unsafe rather than slipping past an allowlist.
Single-replica deployments are fully correct; a fleet is better than it was, not finished. A write landing on one replica leaves the others holding their own copies, and closing that needs the durable purge journal rather than another best-effort fan-out — see 1.8 in ROADMAP.md.
-
The nightly evidence suite keeps its evidence and says when it fails. It uploaded nothing and notified nobody: every number in the capacity model is sourced from a run's console output, which GitHub deletes on its own schedule — a citation with an expiry date — and a suite that runs at 02:47 and reports to no one is a suite nobody reads. Both failures so far were caught only because someone happened to be merging at the time. Each job now captures the k6 summary and PRISM's own logs as a 90-day artifact, and a reporting job opens a labelled issue on any scheduled failure and closes it when the suite is green again. A test pins the artifact retention, the
issues: writepermission, the failure condition and the close path — each of which is useless without the others. -
Release verification now proves the shipped image renders. It checked that the image starts, proxies to the origin and answers
/health— which is precisely what PRISM does when it has no licence, so every release so far was verified against behaviour a plain reverse proxy would also pass. The job now mounts a time-limited CI licence (PRISM_CI_LICENSEsecret), requests the page as a crawler, and asserts bothx-prism-rendered: trueand the presence of content that exists only after JavaScript runs. The origin fixture had to grow up for this: at 92 bytes and titleless it could never satisfy content validation, which is on by default and wants a title, 100 characters of text and 1 KiB of HTML. Missing secret fails the job — an unverifiable release is not releasable. -
A written versioning and deprecation policy (
versioningin the docs): what each version-number component promises, which surfaces are versioned (config schema, admin API, metrics, packages — not internal Rust APIs), and the deprecation contract: nothing disappears within a major, deprecated keys warn at startup naming their replacement (the[detect]section already operates this way), removal only at the next major, anddeny_unknown_fieldsstays strict throughout. -
A shipped ops pack (
deploy/ops/): a Grafana overview dashboard, 12 alerting rules in both plain Prometheus and PrometheusRule CRD form, a runbook per alert, and documented SLOs whose latency objectives anchor to what the nightly evidence suite already enforces. A structural test keeps it honest: every referenced metric must actually be exported, the two rule files must carry the same alert set, and every alert must have a runbook. -
The licence grace window is now visible to metrics.
prism_license_grace_activeis 1 exactly while the window is open — the interval where "renew now" is still actionable.prism_license_validstays 1 through grace (rendering continues), so without this gauge the first alertable signal was rendering having already stopped. -
A published capacity model (
deployment/capacityin the docs): the two-regime cost model (warm hits vs cold renders), the memory formula the binary itself enforces at startup, render-throughput and warmth-budget formulas with a worked example, and crawl-burst levers — every number cited to its source (nightly evidence runs, scale-matrix observations, or code constants) rather than asserted. -
A Chromium update watch with a 7-day adoption SLA. The renderer's pin made builds reproducible and also guaranteed nothing updated it — the one component cargo-deny cannot see. The daily security-audit workflow now compares the pin against Debian's indexes (newest of
bookworm-securityandbookworm, compared with dpkg semantics, failing closed on unreadable indexes) and opens a labelled issue the day a newer chromium ships; SECURITY.md commits to adopting it within 7 days. A structural test keeps the watch's teeth in place. -
Claimed crawlers can be verified with forward-confirmed reverse DNS.
[bot.verify]checks a claimed Googlebot/bingbot/Applebot/Yandex/Baidu identity the way its operator documents (PTR into the operator's domain, forward-confirmed to the same IP), with per-(IP, identity) verdict caching.enabledobserves — verdicts in logs andprism_bot_verification_total{outcome};enforcetreats a proven-spoofed crawler UA as a human — in bot-only mode that means proxied, never buying a render or a Chrome tab. Deliberately asymmetric: only a completed check that positively failed convicts — timeouts and resolver failures areunverifiableand never enforced against, because a degraded resolver must not demote real crawlers to the raw JavaScript shell. Off by default. -
Credential headers are refused as
[cache] varydimensions. A value there becomes part of every cache key and is forwarded to Chrome, sovary = ["cookie"]meant session tokens copied into keys, handed to the renderer, and one cache entry per visitor — an eviction attack by configuration. There is no deployment where that was intended; startup now refuses it. -
Data-handling and disaster-recovery documentation. Both answer procurement rows with facts already true and enforced by tests: the renderer is anonymous, credentials cannot reach it, the cache is memory-only and TTL-bounded (RPO zero — losing every replica loses warmth, not data), and the warmup is the recovery procedure.
-
The licence is re-evaluated at runtime, with a grace window and an in-place reload. It used to be evaluated exactly once, at load: expiry mid-run kept rendering until a restart, and the restart then silently stopped it — two cliffs, both surprises, the second one discovered from an SEO traffic drop. An hourly probe (config:
license.recheck_minutes) now re-verifies the file without touching the live state; a renewal dropped in place is picked up within one interval with no restart; an expiry starts a loud grace window (license.grace_hours, default 72) that is screamed about on every tick before rendering actually stops.POST /license/reloadapplies a renewal immediately, and a reload that fails to verify keeps the previous state — a bad file must never un-license a running fleet. -
Read-only metrics credential.
admin.metrics_bearer_tokenis accepted only forGET /metrics, so Prometheus never holds a token that can purge the fleet; a test proves the scope opens exactly one door. -
Audit events on every mutating admin call — endpoint, caller IP, parameters, entries removed, peer fan-out results — on the
audittracing target, so destructive actions stop being invisible to log review. -
Licence visibility.
/statusreportsversion,licenseandlicensed;/metricsgainsprism_license_validandprism_license_info{status=...}— the first alertable signal for the silent expiry-then-restart bypass. -
Attributed rate-limit and purge metrics.
prism_rate_limited_total{limiter="ip"|"domain"}answers which budget shed a request (the documented CDN 429 incident could not be told apart from domain limiting), andprism_peer_purge_failures_totalis, until durable purge lands, the only sign a purge did not converge. -
A Helm chart (
deploy/helm/prism) mirroring the hardened Kubernetes manifest: same securityContext, probes and resource envelope, a ServiceMonitor wired to the scrape-only metrics token, and a chart that refuses to install without a real configuration rather than inventing one.appVersionis coupled to the crate version by a shipped-config test.
Changed
-
The wire-level contract is asserted positively. The tests checked that a refused response did not carry
max-age=3600, which a missing header or a different lifetime would also satisfy. They now requireno-storeexactly. -
The test fixture builds on a supported runtime, reproducibly. It ran on Node 20 while
react-routerdeclaresengines.node >= 22.22.0— npm warned and carried on, so it worked by accident. It also rannpm installwithout copying the lockfile first, so the fixture the tests exercised was not the one the lockfile describes. -
The wire-level test tier runs in CI. The hurl suite asserts what a socket receives and nothing executed it, which is how three of its assertions came to contradict a deliberate, unit-tested decision — that
Varynames request headers, soX-Prism-Variantmust not appear there — without anyone noticing. Switching it on caught two of the defects above the same day. A test guards that every compose stack undertests/e2eis referenced by CI, because a directory of tests nobody runs reads as coverage while being none. -
scripts/checkis the way to verify the tree. Askingcargo testfor a count of failures answers0when the test binaries did not compile, because a build that produced no results produced no failures either — and piping discards the exit code that said so. Twice that became a confident "0 failures" while nothing had run, hiding a stale duplicate of a whole test module and then three files that no longer matched a struct they construct. The script asserts that tests ran and passed, refuses a run below a floor that would indicate a target silently not being built, and keeps cargo's exit status. -
The documentation site's canonical URL points at the host that serves it. It named
docs.trident-prism.com, which 301s todevdocs.trident-prism.com— so every page told search engines its canonical address was one redirect away from itself. Awkward anywhere, specific here: PRISM exists to control how crawlers see a page. -
The render/bypass decision is a function, and its reason is a type. The choice was an inline if-else chain in the request handler with a string literal at each branch. The strings were already stable and already logged; what they could not do is be enumerated, so nothing could assert that the decision covers its own input space and nothing downstream could act on a reason without matching text.
It is now
decide_request(RequestFacts) -> (RequestDecision, DecisionReason)— noself, no request, nothing but the eight facts already computed by the time the choice is made. The logged strings are unchanged, so existing queries and dashboards keep working.What that buys is testing the whole space rather than the paths a test happens to walk: all 256 combinations of the eight facts across both modes, asserting properties instead of a table. Nothing renders without a licence whatever the other seven say. Bot-only never renders for a human — the property the privacy posture rests on. A render implies every condition was satisfied. Every reason is reachable, so no branch is dead code wearing a name. Making the licence check reachable-past fails the first of those by name.
-
Only the statuses a cache is allowed to store are stored. The rule that kept 401, 403, 407 and 429 out of the cache unless the origin said otherwise is now the rule RFC 9111 §3 actually states, asked as "is this on the permitted list" rather than "is this one of four".
The case that showed it was the temporary redirect. A 302, 303 or 307 was kept for the configured hour — so a page moved aside for maintenance, an experiment or a geography stayed moved for crawlers long after it came back. A redirect that calls itself temporary is the one thing a cache should not pin. Their permanent counterparts, 301 and 308, are on the list and still store, as do 200, 203, 204, 206, 300, 404, 405, 410, 414 and 501.
Server errors are exempt, and the exemption is where it belongs. A 5xx answers to
cache.error_ttl_secs, which is the operator's explicit permission, so it does not pass through the §3 gate — applying it there would refuse every server error and switch off error caching entirely, the deliberate behaviour that shields a struggling origin. That regression was caught by an existing test rather than by review. -
One cache decision, computed once. Storage lifetime was decided in the render pipeline while the client-facing
Cache-Controlwas built separately in the proxy, and the storage rule itself was written twice — once for the foreground render, once for the background refresh. Four defects fixed earlier in this series were each an instance of the same shape: two places deciding, so they could differ. A 5xx kept for thirty seconds and announced for an hour; statuses stored that should not have been; an origin policy with no number left as a bare word; an entry whose clock started after the render that produced it.EffectiveCachePolicynow decides all of it at once — whether to store, why not when it will not, for how long, from what age, and whether the grace window may serve it. Both render paths call it; the duplicated rule is gone.The refusal reason is the new capability. A page that would not cache used to say so with a header value and nothing more, and answering why on a live deployment has taken days. There are four reasons, each a different conversation and each carrying both a stable token for metrics and a sentence for a person: the origin forbade shared caching; the origin granted a lifetime of zero, which is a permission without a duration rather than a prohibition; the status needs the origin's explicit permission; error caching is switched off.
What guards it is an invariant test rather than a list of cases: across every status, origin directive and error-TTL combination, whatever is announced must equal whatever is kept. Reintroducing the original 5xx defect fails it with "kept for 30s and announced for 3600s".
-
A cached entry no longer starts life younger than it is — RFC 9111 §4.2.3. PRISM read the origin's
Ageheader, ignoredDateentirely, and started the entry's clock at the moment of storage. Both halves were wrong, and the second one especially for this product: the document arrives from the origin and then spends seconds inside Chrome before anything is stored. On the shipped 30-second error TTL a five-second render was a sixth of the lifetime, granted for free because nothing counted it.The corrected initial age is now the larger of two estimates, as the RFC requires: what
Dateimplies, and theAgeheader plus the time the fetch and render actually took. An origin behind its own cache reportsAge; an origin in front of nothing reports onlyDate; a response that passed through both is credited with whichever is more honest about its age. ADateahead of our clock is treated as skew rather than as a response from the future, and an unreadable one as absent — the same reasoningAgealready carried. -
A cache policy without a number is completed rather than ignored. An origin sending
Cache-Control: publichas stated a policy and no lifetime. PRISM read the header's mere presence as a decision, suppressed its own synthesis, and stored the entry for its configured TTL anyway — an hour, by default. Downstream was left to derive freshness fromETagandLast-Modified, which rendered responses deliberately do not carry because the body was transformed, so it had nothing to work from at all. PRISM now keeps the origin's own words and appends the lifetime it is actually using:publicbecomespublic, max-age=3600, s-maxage=3600. A header that already carriesmax-ageors-maxageis still left untouched, and one that forbids storing —no-store,private,no-cache— is never given a lifetime, because PRISM is not keeping it either. -
A momentary refusal from the origin is no longer given the route's lifetime. 401, 403, 407 and 429 were stored under the configured TTL whenever the origin sent no directive of its own. None of them describes the page: they describe this request. A 429 raised for one second by an origin's rate limiter was kept for the configured hour and served to everyone after it, turning the origin's own protection into an outage that outlasted what it was protecting against. RFC 9111 §3 does not list these among the statuses a cache may store heuristically, and PRISM no longer does either — unless the origin states a lifetime for one, in which case it is honoured and capped as usual.
skip_authenticatedalready keeps credentialed requests out of the cache; this covers the unauthenticated request that happens to receive one of these. Every other status is untouched, which has its own test. -
PRISM's own error responses carry
no-store. Its 429, 502, 503 and 504 are statements about the current instant — a limiter that is full now, an origin unreachable now. A CDN that stored one would go on refusing clients PRISM would already serve, and PRISM would never see those requests to know it was happening. -
A cached error is no longer advertised for longer than it is kept. PRISM synthesises
Cache-Controlwhen the origin sent none, and that synthesis read the route's lifetime without looking at the response status. The render path meanwhile stores a server error undererror_ttl_secs. With the shipped defaults the two disagreed by a factor of 120: a rendered 5xx was held for 30 seconds and announced aspublic, max-age=3600, s-maxage=3600, so a CDN or browser in front of PRISM kept an error page for an hour — long after the origin had recovered and long after PRISM itself had dropped it. The announced lifetime now matches the one actually used, and a 5xx that will not be stored at all (error_ttl_secs = 0) is announcedno-storerather than offered downstream. Responses that are not server errors are unaffected.This changes what downstream caches do with your error pages. If you relied on the old behaviour to absorb origin outages at the CDN, raise
cache.error_ttl_secs— that setting now controls both ends, which is what its name always implied.The decision moved out of the request handler into
synthesized_cache_control, which had no test coverage of any kind before this change: the name appeared only insrc/server.rsandsrc/config.rs. It now has seven, including the boundaries of the 5xx range and the case where the origin stated its own policy. -
The memory guard is calibrated to the heavier of the two workloads we measure.
CHROME_BYTES_PER_TABcame from the pool soak (375 MiB/tab on ordinary pages) while the nightly scale matrix settles the same 4-tab pool at 2.4–2.6 GiB on adversarial ones — so the guard stayed silent exactly when pages got heavy, and the shipped 8-tab default appeared to fit 4 GiB. Now 525 MiB/tab, derived from that observation, with a test that fails if the default pool ever silently fits a 4 GiB limit again. -
The licence is commercial everywhere, and the crate can no longer be published.
LICENSEwas MIT,Cargo.tomlsaidlicense = "MIT"and the README footer offered the project under it — while the product documentation and the binary itself say PRISM is commercial and refuses to render without a signed licence file. The repository was granting the rights we invoice for.LICENSEnow states commercial terms conditioned on that licence file,Cargo.tomlcarrieslicense-filepluspublish = false(crates.io never deletes a published version, so onecargo publishwould have been permanent and irreversible), andcargo-denyskips unpublishable crates while still checking every dependency. Package metadata declared MIT too —dpkg -s prismandrpm -qi prismprinted it on every release so far — and pointed its homepage at the internal source repository, which answers 404 to every customer who follows it; both now say what is true. Since the software ships without source, the terms travel with it: packages and the container image carryLICENSE. A structural test keeps every surface in step. -
GitHub Releases are retired; provenance moved where customers can reach it. The repository is internal, so release assets answered 404 to everyone a release is for — installation is the package repositories and the Docker registry, and always was. Each release now publishes its CycloneDX SBOM and SHA-256 checksum manifest at
pkg.trident-cache.com/prism/provenance/<tag>/, staged, verified and promoted through the same fail-closed chain as the packages; structural tests pin the ordering and the production assertions. The publish phase drops itscontents: writetoken (the Release was its only writer), the broken-by-designdist/install.shis gone, andCHANGELOG.mdis the release record. -
The origin's cache directives now decide what PRISM keeps. Only
no-storeandprivatewere read; everything else was passed to the client and ignored on the way in. A page markedno-cachewas stored, reused without validation, and served alongside the very header saying that was not allowed — stale prices and stale availability, advertised as fresh.One limitation decides most of this: PRISM cannot revalidate. There is no conditional request back to the origin before a stored render is reused, so a directive asking for validation before reuse can only be honoured by not keeping the entry. That now covers
no-cache,max-age=0,s-maxage=0, andVary: *, which must never be reused for another request.Where the origin does grant freshness it caps what PRISM keeps rather than being overruled by it:
s-maxagefirst, since it exists to address shared caches, thenmax-age. A page the origin called fresh for a minute is no longer served for the hour PRISM was configured for.Grace is part of that arithmetic rather than an exception to it. PRISM serves an entry for
grace_period_secspast its TTL while re-rendering, which is exactly whatmust-revalidateforbids, so for those the TTL is pulled back until the whole window fits inside the granted freshness. If nothing is left, nothing is stored.If your hit rate drops, look at what your origin sends on those routes. An origin that means pages to be cached should say so;
s-maxageis the directive aimed at PRISM specifically. -
A shared cache may no longer stretch what the origin granted.
s-maxagecarries the semantics ofproxy-revalidate(RFC 9111 §5.2.2.10), so grace — which serves an entry past its TTL while a re-render runs — must not apply to it, and did.must-revalidatewith no stated lifetime was ignored on the same path. The entry now records the origin's refusal and grace skips it, which also fixed the first attempt at this: subtracting the grace window from the TTL refused anything whose freshness was shorter than the grace period, so ans-maxage=60page went uncached under a five-minute grace. -
Responses now carry
Age, and inherit the one they arrived with. RFC 9111 §4 requires a cache to say how old a reused response is. PRISM said nothing, so a downstream CDN started its own freshness clock from zero and the same seconds were counted on every hop. IncomingAgeis now subtracted from remaining freshness too:max-age=60withAge: 50has ten seconds left, not sixty, andAge: 90arrived stale. -
A render the origin refused to a shared cache is no longer handed to the requests waiting on it. Coalescing broadcast it to every concurrent waiter, reaching them as
x-prism-cache: HIT— a hit on something never cached. Not storing it did not make it less shared.This costs renders. Four concurrent requests for a
no-cacheURL are now four renders rather than one, each charged torate_limit_per_ip— three by default, and one address for all traffic behind a CDN unlesstrusted_proxiesis set. Size that budget before upgrading if your origin sendsno-cacheunder crawl. -
An authenticated shopper is sent to the origin instead of rendered. In render-all, Chrome renders with no session — no cart, no account, no prices that depend on who is asking.
skip_authenticatedkept that result out of the cache, so one shopper's page never reached another, but the shopper who asked still received it: an empty cart on their own cart page.A request carrying an Authorization header or any cookie is now proxied to the origin. Anonymous visitors on the same route still render. Bot-only is unchanged, and setting
skip_authenticated = falserestores the old behaviour.
Fixed
-
Sitemap warmup no longer refuses the origin it was pointed at. The private-address resolver added in 1.7.1 enforced regardless of
block_private_cidrs = falseand of the configured-origin exemption, sohttp://spa:3000/sitemap.xmlfailed whereverspais a container name on a private range — most compose and Kubernetes deployments, and the shape the documentation uses. Literal-IP testing never showed it, because a literal is not a name. The URL check and the DNS check now ask the same question. A host that is not the origin still may not answer with a private address. -
A coalescing waiter registers atomically. When a leader's client disconnected, every waiter woke at once, each found nothing in flight, and each registered — several rendered the same URL while the rest were told the tab pool was full, a 503 for a page that renders perfectly well. The recovery path uses the same atomic operation the normal path always did, and an in-flight entry is now removed only by the registration that created it: a disconnected leader's guard no longer deletes the waiter that replaced it.
-
A recycle drains the tabs that exist. It waited for the configured tab count from a semaphore that can hold fewer — a launch where one tab failed to open, or a forgotten permit, shrinks it. A 7-of-8 pool could never drain: it waited out the whole 35-second budget, logged "renders still in flight" when there were none, then restarted underneath the healthy renders it never waited for. Those died as crashes and fed the circuit breaker, so one tab lost at boot degraded every recycle afterwards.
-
A completed render no longer walks the process table. The browser-memory gauge was refreshed inside
release(), which the request awaits before it can answer, so every render paid for aread_dirof/procplus two file reads per process — the host's whole table under the shipped systemd unit, which sets noProtectProc. The five-second sampler already keeps the gauge current. -
Admin sitemap warmup resolves through the private-address guard. The URL check looked at the name, and the fetch that followed resolved that name again — nothing required the two answers to agree, so a host that answers public once and private on the next query reached an internal address from inside the render host. Checking twice does not close that; there is always another lookup after the last check. The client now resolves through the same classifier the render path uses, and refuses a name that fails, times out, returns nothing, or includes a private address.
-
A release proves the commit it publishes passed its tests.
release.ymlrefused a tag that was not onmain, which says the code was reviewed and merged — not that it worked. The automatic trigger fires only after the required jobs succeed, but a manualworkflow_dispatchreached the publish phases without any of that, and every release so far has been a manual dispatch: green by habit rather than by construction. A published version number cannot be withdrawn here, so the check belongs before the publish. -
A retry cannot double the wait the caller was promised. Each render attempt carries its own global budget and the retry started a second one from zero, so with the shipped ten-second timeout a request could wait two thirty-second budgets before falling back — in render-all, that is a person sitting through both. The retry now gets what is left of one total budget rather than a fresh one of its own.
-
Repeated origin headers no longer corrupt cache policy or panic a connection. CDP joins repeated fields with newlines.
privatefollowed bypublic, max-age=3600was consequently stored as public; other combinations reached an infallible header conversion and closed the connection on every MISS and HIT. Values are split back into repeated fields, the effective policy combines everyCache-Control, and the final conversion fails safe tono-store. -
An iframe can no longer become the page's status and headers. Document response events are restricted to the main frame, so an embedded 404 with
no-storeandnoindexcannot turn a cacheable 200 page into that response. -
Render-all keeps human traffic available when Chrome has no capacity. A full tab queue or open circuit now sends crawlers a cache-safe 503 with
Retry-After, while JS-capable human clients receive the origin response as documented. Tab acquisition and pre-navigation CDP setup are also inside the global render lifecycle budget, so a wedged browser cannot bypass that cap. -
Tracking parameters cannot poison their shared cache entry. Chrome now navigates the same normalised URL used by the cache key.
-
Turning synthesis off does not turn safety off. Cache-disabled and authenticated cache-bypass responses with no origin policy still emit
no-store; a downstream cache is never invited to keep what PRISM kept nowhere. -
Encoded URLs can be inspected, invalidated and purged. Administrative and unsafe-write paths now escape canonical URL prefixes exactly as cache-key construction does.
-
Server errors stop at
cache.error_ttl_secs. The ordinary page grace window no longer extends a 5xx beyond its advertised error lifetime. -
Page-supplied JavaScript statuses are range checked. Values outside 200–599 no longer become PRISM's final HTTP status.
-
Origin
Datereaches corrected-age calculation in production. The tested Date branch is no longer cut off by the preserved-header filter. -
/explainrecognises variant-only entries, authenticated storage bypasses, effective preserved origin policy, and encoded cache keys. -
Malformed sitemap XML is an error, not an empty or partial success, and a sitemap index cannot pivot recursive fetches outside its parent authority.
-
Turning the cache off now turns it off downstream too. With
cache.enabled = falsePRISM stored nothing and went on announcingpublic, max-age=3600, so an operator who disabled caching had a CDN in front keeping rendered pages for an hour — with nothing in PRISM's own configuration to explain where the stale copies came from. The documented meaning of the setting is that caching is off, and downstream is part of "off". An origin that stated its own policy still keeps it: disabling PRISM's cache is not an instruction to countermand the origin. -
A credentialed request is no longer offered to a shared cache.
cache.skip_authenticatedsends a cookie-bearing request past PRISM's cache, because the body answers that requester. PRISM kept none of it and announcedpublic, max-age=3600anyway — the same failureskip_authenticatedexists to prevent, moved one hop downstream. That path also reported the origin's rawAgerather than the corrected one, and now goes through the same decision as every other. -
Sitemap numeric references and CDATA are read.
&and&are the decimal and hexadecimal spellings of&, and arrived as literal text — so warmup fetched?x=1&y=2, a URL that does not exist, and reported success for it. A<loc>wrapped in<![CDATA[...]]>, which is the other legal way to write an unescaped&and what several generators emit, produced no URL at all: in the warmup report a silently dropped page and a page nobody listed look identical. -
A refusal explains itself in terms of the status it was asked about.
/explaingave every status outside RFC 9111 §3's list the sentence "401, 403, 407 and 429 describe the request rather than the page" — including for a 302, sending the reader after an authentication problem that does not exist. -
A sitemap URL was truncated at its first escaped character. XML requires
&to be escaped, so?a=1&b=2is what a correct sitemap generator writes — and the parser stopped at the ampersand, warminghttps://example.com/search?a=1instead. Every URL with more than one query parameter was affected, and silently: warmup reported success for a page it had never fetched.The cause is that a
<loc>does not arrive as one event. quick-xml splits the content at every entity reference —Text("?a=1"),GeneralRef("amp"),Text("b=2")— and the parser took the first fragment and stopped reading. It now accumulates until the closing tag, resolving the five predefined entities and passing anything else through as written, since a URL that 404s is visible in the warmup report while a dropped entry looks exactly like a page nobody listed.Present since the sitemap parser was written; found while migrating to quick-xml 0.42, by writing a test for behaviour the migration touched.
-
What PRISM refuses to store, it no longer invites a CDN to store. 1.6.0 stopped keeping a 302, 401, 403, 407 or 429 without the origin's permission — and went on announcing
public, max-age=3600on the way out. So PRISM declined to hold a 403 for a second while inviting Cloudflare, Fastly or Varnish to hold it for an hour. For a renderer whose job is how crawlers see a site, that is an outage with our name on it: a 403 pinned at the edge is a page Googlebot cannot reach until the TTL expires. It was also worse than before 1.6.0, which at least kept the two halves consistent.The announcement is now derived from the storage decision instead of computed beside it, and carried on
RenderResultand the cache entry rather than recomputed by the proxy — so a hit and the miss that produced it cannot answer the same question differently. When storage is refused and the origin said nothing, the response carriesno-store; when the origin said something, it is left exactly as sent. -
A URL whose query parameters are all stripped now shares the bare URL's cache key.
normalize_cache_keyreturned its input verbatim when there was no query to strip, skipping the normalisation every other path applied. With an origin configured ashttp://shop:80— an explicit default port, which is legal — that produced two keys for one page, so a crawler arriving with?utm_source=...never hit the entry warmed for the bare URL and paid for a full render every time. Silent: the cache appears to work, the hit rate is simply lower than it should be. Present since the function was written. -
The age reported on a miss is the corrected one. A ten-second render was advertised downstream as zero seconds old, so a hit and the miss that produced it reported different ages for one document. Coalesced waiters had the same defect.
-
The shipped memory ceilings no longer make the binary warn about itself. 1.5.0 recalibrated Chrome's per-tab budget from 375 to 525 MiB against the nightly scale matrix, which moved the startup guard's estimate for the default eight-tab pool from 3400 to 4600 MiB. The ceilings shipped alongside it did not move.
dist/prism.service— and therefore every deb, rpm and apk of 1.5.1 — keptMemoryMax=4G, which is 4096 MiB, so a fresh install warned at startup that Chrome would be OOM-killed under load and that a restart policy would turn that into a restart loop. The unit is nowMemoryMax=6G; the Docker and Kubernetes ceilings were already 6, and the Kubernetes and Helm requests move from 4Gi to 5Gi so the scheduler reserves above the working set rather than below it. A test now parsesMemoryMaxout of the shipped unit, reads the pool size a default config actually produces, and runs the real guard against them — string-matching the shipped files against each other is what let this through, since they agreed with one another while both were wrong. -
The leak gate asks whether Chrome's RSS plateaus, not whether it grew.
repeated_large_renders_do_not_leak_resourcescompared a baseline against a final reading, and both attempts to make that comparison meaningful failed. Warming up harder made it worse, because the warm-up is itself enough to trigger a recycle. Skipping the comparison when the recycle counter moved during the run still allowed a recycle just before it: the nightly reported 1178 MiB growing to 2517 MiB while every other case in the same run sat between 2397 and 2546, which is a browser seconds old reaching the ordinary steady state. Two readings cannot separate warm-up from a leak, because both grow; what separates them is that warm-up flattens. RSS is now sampled after each of six rounds, restricted to the longest trailing run of rounds sharing a recycle counter, and the final round-to-round step must be under 256 MiB. The selection is a pure function with tests that run without Docker, including one built from the failing run's real series. -
Nightly results are reported when the suite is re-run by hand. The reporter carried
github.event_name == 'schedule'. The workflow has exactly two triggers, so that excluded the manual re-dispatch: a re-run that went green could not close the issue a failure had opened, and a re-run that failed could not update it. It looked like it worked — an issue opened on a scheduled failure and closed the next morning — but that close was done by hand while the job was skipped. Guarded by a test that checks only the jobs callinggithub.rest.issues, since selecting a soak tier from a dispatch input is a legitimate use ofevent_namein the same file. -
The Chromium watch can close the urgent issue it opens. It opens two titles — "update available" and "no longer installable" — but both the retitle and the close lookups matched only the first, so the urgent one could be opened once and never found again. Both lookups now recognise both, with a test that counts the matchers on executable lines: the first version of that test counted the phrase, which also appears in the comments explaining it, and passed while both lookups were still wrong.
-
The chart pins its rollout strategy instead of taking Kubernetes' default. 25%/25% removes a quarter of the render capacity mid-rollout and brings the replacements back with a cold cache — a default nobody chose, while the manifest has pinned
maxUnavailable: 0, maxSurge: 1all along. values.yaml has claimed since it was written that its "defaults mirror deploy/k8s"; the only thing checking that was the version string. A coupling test now covers the policy's excluded ranges, the strategy on both sides, and the NOTES warning, and CI asserts them in the rendered output. -
acquirenow honours the one deadline it was given. It set a deadline, then looped — fail fast, wait out a recycle, try again — while each attempt started a freshACQUIRE_TIMEOUTof its own, so a call could take twice the bound its caller was promised. That matters beyond the wait: the render pipeline's coalescing deadline is computed from a model that says one timeout, so waiters could give up on a leader that was still inside its legitimate budget and turn a slow recycle into 503s for exactly the requests coalescing exists to absorb. Each attempt is now bounded by what remains, via a pureattempt_budgetin the style of the pool's other decision functions, with a test that fails against the old behaviour. -
The five-minute guide no longer hides the licence. It never mentioned licensing — zero occurrences — while instructing the reader to request a page as Googlebot and expect "fully rendered HTML with all JS-generated content". A fresh install has no licence, so PRISM proxies and renders nothing, and the very first thing every evaluator tried returned the raw shell from the page whose job is their first five minutes. The guide now lists the licence as a prerequisite, mounts it in the
docker run, says plainly what an unlicensed install looks like (indistinguishable from a working one except for the rendering), and points at/status→licensedas the first thing to check. A test keeps it that way. -
The Chromium watch could not open the issue it exists to open. The patch that taught it to distinguish "a newer version exists" from "the pin is no longer installable" was lost mid-edit: the issue body referenced an
installablevariable the compare step never defined, so the stale path would have thrown a ReferenceError and opened nothing. Both halves are now present and the file was read back to prove it. -
Ten documentation claims that did not match the code.
admin-api.mdreported a warmup status (completed) the enum never emits (complete) and implied the mutating endpoints are open when no token is set — PRISM generates and logs one, and refuses a non-loopback bind without either a token or an explicit opt-out. README named a metric that does not exist (prism_render_duration_sum_ms; the export isprism_render_duration_ms_sum) and a test count from an earlier era. SECURITY.md recommendedMemoryMax=2Gagainst README's 4 GiB floor, and told Docker operators to setblock_private_cidrs = false— the change a shipped-config test refuses, and unnecessary, since the configured origin is already exempt.versioning.mdpromised support for two minor series where SECURITY.md's table says the current one.capacity.mdcalled the manifest's 4 Gi a limit (it is the request; the limit is 6 Gi) and described the shipped pool as the code default (4 tabs versus 8), and cited two measurements to a k6 run that reports neither — one is derived from PRISM's own request logs and now says so.metrics.mdcalled itself a complete reference while omitting 18 of 42 families, and shipped competing alert definitions reusing the names indeploy/opswith different thresholds; the families are documented and the alerts now point at the tested set. Two tests keep the reference complete and single-sourced. -
Documented configuration that refused to start. The CDN recipe in README and cdn.md recommended
vary = ["x-prism-variant"]— rejected at startup by the credential-vary validation 1.4.2 itself added, and contradicting the warning two paragraphs above it, which correctly saysX-Prism-Variantis a response header. The edge header is nowx-device-class, a name the edge owns, with the collision explained. troubleshooting.md's two remedies for HTTP 429 putrate_limit_per_ipunder[server]andrate_limit_per_domainunder[render]; both live under[security], both are per second not per minute, and an operator copying either during an incident got a process that would not come up. The React and Angular presets dropped the requiredselector:prefix. A new test parses every TOML block in README, SECURITY.md and the documentation site through the real deserializer andvalidate()— 103 blocks, all starting — so the whole class stays closed rather than these four instances. -
Both shipped Kubernetes paths now produce a deployment that renders. The raw manifest — which the chart's README calls the reference — carried no licence volume, no mount and no
[license]section, so it could not render by construction; the Helm chart mounted the Secret as a directory at/etc/prism/license, leaving the file at/etc/prism/license/license.licwhile every document says/etc/prism/license.lic. Both failures were silent: pods go Ready, metrics look healthy, and crawlers get the JavaScript shell. The manifest now mounts a requiredprism-licenseSecret and points its config at it, the chart usessubPathto land the file exactly where the docs promise, and an unlicensed install says so in NOTES rather than looking like a working one. -
CI renders the deploy artefacts instead of only reading them. A new job runs
helm lint, templates every mode the chart supports (licensed, unlicensed, autoscaling + PDB + ServiceMonitor), asserts the licence lands on the documented path, checks the chart still refuses a missing configuration, and parses the raw manifests as Kubernetes objects. The chart is the path we recommend and nothing had ever rendered it. -
Chromium pin moved to 151.0.7922.173, which is a security update, and the watch now knows the difference between "behind" and "unbuildable". Debian's indexes carry only the current version of a package, so when 151.0.7922.173 superseded .137 the pinned version stopped existing and every image build began failing — CI, e2e and releases alike, within hours of the daily watch reporting the pin as current. The watch now reports two distinct conditions: a newer version exists (adopt within the 7-day SLA) and the pin is absent from every index (broken builds, fix now), with the issue title and body saying which.
-
Cache keys are injective again — a collision was an unauthenticated poisoning primitive.
normalize_cache_keypercent-decoded the query withquery_pairs()and rebuilt it withformat!("{k}={v}"), so the separators came back as structure:?a=b%26c%3Dd(one parameter whose value contains&and=) and?a=b&c=d(two parameters) produced the same key. Because the render fetches the raw URL and stores the result under that key, one GET for the encoded form parked the attacker's page on the victim's entry — no credentials, no headers, no timing. The query is now re-encoded when the key is rebuilt, and three tests pin it: the collision, the same trick aimed at thestrip_query_paramsdenylist, and idempotency (a key fed back through normalisation must not move, or purges stop matching what is stored). Found by the 1.4.2 audit, which reproduced it by execution rather than by reading. -
A grace window now closes even over a deleted licence file.
reloadrefuses to mutate state unless the file verifies — right for a bad renewal, wrong as the only tool for grace expiry: a licence file deleted or corrupted after expiry left the live state licensed forever, contradicting the documented "rendering stops when the window closes". The watcher now sheds the licence explicitly when reload cannot (trident-license gainedmark_expired), with the behaviour pinned by tests on both sides. -
Documentation claims brought in line with what the code enforces. An adversarial docs-vs-code audit found seven statements without backing:
--versiondoes not report licence state (the startup log does), packages shiplicense.pathcommented out rather than defaulted, the auto-generated admin token is deliberately printed once at startup (the docs claimed "never logged"), and disaster-recovery quoted timing numbers and an 8-tab reference pool that nothing measures or runs (the CI stacks use 4 tabs). Docs now state the enforced numbers and the real defaults. -
The k6 scenarios' failure accounting matches their own semantics. The spoofed-bot scenario exists to provoke 429s, yet hung its threshold on k6's built-in
http_req_failed, which counts every 4xx/5xx — it could never pass (first execution: 98.9% "failed" while 98.3% of checks succeeded). Three sibling scenarios carried the same latent mismatch and passed only because their rates happened to stay under threshold. Every scenario now excuses exactly the statuses its own checks accept (per request where bots and humans differ — a rate-limited human still counts as a failure), and the spoofed-bot scenario's ownfailed_requestsmetric finally has a threshold. -
The k6 load stack can actually launch Chrome. The nightly evidence suite's first execution failed all five k6 jobs: the compose lacked the sandbox settings (
SYS_ADMIN,apparmor=unconfined, 1g/dev/shm) every e2e compose already carried, so Chrome could not create user namespaces. Mirrored verbatim and proven locally end to end before shipping. -
Single-URL and pattern purges removed nothing in viewport mode. Cache keys carry the device class as
::mobile/::desktop; both purge sweeps anchored on|alone, so a price-change purge matched zero entries cluster-wide, reported success, and stale prices served to Googlebot until TTL. The tests covering the sweeps fabricated a|mobilesuffix that key construction never produces — they now build every key through CacheKey itself, and matching derives from the sameDeviceClass::ALLthe keys do. -
Warmup filled a cache its intended reader never looked at. With viewports enabled, warmup and admin
/renderrendered only the desktop variant, so Googlebot Smartphone — the crawler mobile-first indexing sends — missed on 100% of warmed URLs. What "warmed" means now comes from the pipeline that owns the keying decision, and both callers iterate it. -
Coalesced waiters gave up before their leader legitimately finished. The waiter deadline was
timeout+5s; a leader may wait out a 60s tab acquire and spend the 3× global render budget, twice. The deadline is now derived from those same constants, so the two cannot drift. A leader that dies still wakes its waiters immediately. -
Wedged-Chrome recovery waited behind its own readers. Restart takes the browser write lock while read guards held CDP calls bounded only by the transport timeout (sized above a whole render), so recovery took 1–2+ minutes of full-instance 503s after a ten-second detection. Every CDP call made under a read guard is now bounded at 5s — control-plane calls finish in milliseconds on a healthy browser. Replacement-tab creation retries with a fresh guard between attempts before a slot is ever given up: one timeout cannot tell a wedged Chrome from a healthy browser on a saturated host, and treating it as a permanent loss would have let load drain the pool.
-
HEAD to a large resource returned 502. The Content-Length early-reject fired for every method, and a HEAD's declared length describes a GET body the origin never sends. Error responses in the proxy path also no longer build bodies for HEAD; the method-blind error helper is deleted.
-
Malformed
trusted_proxiesentries were silently dropped, and a dropped entry is indistinguishable from an absent one — a whitespace typo reproduced the documented site-wide CDN 429. Broken entries are now named in a startup warning (a startup error from 1.5.0), the parser rejects impossible prefixes, and IPv4 entries now match the::ffff:mapped peers a dual-stack listener reports — previously "trust nothing" on[::]binds. -
Startup warnings finally reach a logger. validate() warned before the tracing subscriber existed, so the memory-budget OOM warning and the
[detect]notice had never printed once. They are returned to main and logged after logging initializes. -
The documented installer failed checksum verification on every release ever made — it saved the artifact under a renamed file and verified against the published name. Files keep their names end to end now.
-
deploy/k8s/deployment.yamlshipped v1.3.3 after 1.4.0 released. The manifest now carries the crate version and a test couples the two: bumping either alone fails the build. -
The response direction strips the full RFC 9110 §7.6.1 hop-by-hop list, matching the request side.
-
A bypassed page now says what it varies on. In bot-only mode a URL has two representations: what a crawler is rendered and what a person is proxied. Only the rendered one declared the dimensions telling them apart, so a downstream cache could store the human copy against the bare URL and answer a crawler from it — a request PRISM never sees, and so cannot correct. Both representations now declare the same dimensions, which is what RFC 9111 §4.1 matches on. Routes that never render claim no variance, so edge caches are not fragmented for nothing.
-
Rendering no longer strips the origin's security headers. The preserved list carried three and dropped the rest, so a crawler fetching the rendered copy got no
nosniff, no referrer policy and none of the cross-origin isolation headers — while a human on the bypass path, on the same URL, got all of them. -
Every way out of the render path declares the same variance. Not being a bot was one of several: a response that was not HTML, and a render that failed and fell back, both went out declaring nothing.
-
Cache directives are read conservatively. The last duplicate value won, so a second
max-agecould widen what the first granted, and an unparseable value was dropped silently — which read as "the origin said nothing" and handed the response PRISM's full TTL. RFC 9111 §4.2.1 asks for the first occurrence, or for the response to be treated as stale. -
A background refresh can revoke what it was refreshing. One that came back
no-storeonly declined to write the new copy; the old one stayed inside its grace window and kept being served. -
The 5xx error budget is bounded by the origin.
error_ttl_secssays how long PRISM will hold an error, not how long it may outlast what the origin granted. -
Headers named by
Connectionare removed before forwarding. RFC 9110 §7.6.1 requires an intermediary to drop them: they describe the hop ending here. PRISM removed a fixed list and never read what the sender actually named, in either direction. -
TTL rules are matched against the path without its query. A request carrying a query got two TTLs — the render pipeline asked with the query attached, the
Cache-Controlheader was built without it — so an entry could be stored under one and advertised with the other for the rest of its life. -
purge_allno longer loses track of bytes. It cleared the entry map, released the lock, and only then zeroed the byte counter. A store landing in those two instructions added bytes to a total about to be overwritten, and its entry stayed in the map with nothing accounting for it. Eviction decides against that counter, so entries it cannot see never create eviction pressure.
Security
-
Every Chrome option PRISM supplies now reaches Chrome. chromiumoxide adds the two leading dashes itself; PRISM supplied them too, producing four-dash options Chrome ignored. This included the internal-host resolver rules and the requested new headless mode. A Chrome-backed regression test now reads the real child process command line rather than the builder's debug output.
-
Percent-encoded excluded routes stay excluded.
/%61dmin/secretcan no longer evade the default/admin/**rule and reach the renderer. Invalid escapes, control characters and dot segments fail closed. -
Forwarding identity and rewrite headers are allowlisted by construction. Vendor aliases supplied by the client can no longer override identity or routing in an origin that trusts them.
-
The deployment guide now states the actual browser boundary: CDP Fetch interception is defence in depth, while network/DNS egress controls are required to cover WebSocket, WebRTC, popup and out-of-process-frame traffic.
-
The sitemap size limit now prevents an allocation instead of reporting one. Warmup downloaded a sitemap in full and then compared its length against the 50 MB ceiling — the comment beside the check said as much, "post download size check". An origin that omits
Content-Lengthand streams five gigabytes would have PRISM hold all five before the limit had anything to say. The body is now read as it arrives and refused at the first chunk that crosses the ceiling, so the buffer never exceeds it by more than the chunk in hand. TheContent-Lengthcheck stays as a courtesy: it only helps against an origin that declares its size honestly, which is not the one this guards against. -
The two open
image-sizeadvisories are assessed rather than left unexplained. GHSA-w3rx-r6r6-pgpr and GHSA-5p2g-fcmc-qvqq are infinite loops in its ICNS, JXL and HEIF parsers. It arrives transitively through@docusaurus/mdx-loader, runs only while building the documentation site, and appears in no binary, image or package PRISM ships; no fixed version exists upstream. Those parsers are entered only for those three formats and the documentation contains none, so a test now fails the build if such a file is ever added. Recorded in SECURITY.md, because an open high-severity alert with no written assessment reads the same as one nobody looked at. -
The Helm chart ships the egress policy the raw manifest has had since 1.3.0. It shipped none, so the path the README recommends was the less defended one: no network backstop for the SSRF case PRISM's own filter cannot fully cover — a public hostname that resolves to a private address, checked for the top-level navigation but not for every subresource. On by default, with link-local (cloud metadata) and CGNAT excluded like the manifest. The origin rule has no default, because a guessed selector renders a policy that looks complete and blocks the one destination the product exists to reach; leaving it unset omits the rule and NOTES says so.
-
A URL can no longer forge the vary section of its own cache key. Vary names and values were escaped; the URL in front of them was not — so a request for
/p|vary:accept-language=deand a request for/punder anaccept-languagevary dimension produced the same key, and whichever arrived second was served the other's HTML. Unauthenticated, and reachable on any deployment with[cache] varyconfigured. The URL portion now escapes|and%(and only those:;and=cannot be misread, since vary begins only after a literal|vary:, and escaping them would rewrite every ordinary query string in every key).purge_patternescapes the operator's pattern the same way, so patterns keep matching and globs are untouched. -
The service-account token is no longer mounted into the pod that runs a browser. Neither Kubernetes path set
automountServiceAccountToken: false, so a renderer compromise handed the attacker this pod's identity with the API server — for nothing, because PRISM never calls the Kubernetes API. Off in the manifest, off by default in the chart, guarded by a test. -
A shipped placeholder credential is refused at startup.
deploy/k8s/config.tomlcarriesbearer_token = "CHANGE-ME-BEFORE-DEPLOYING"so the field is visible where an operator looks — but applying that tree unedited produced a running admin API whose token is printed in a public example, granting purge, warmup, render and licence reload.validate()now refuses any obvious placeholder and names the field. Omitting the key instead would be worse under Kubernetes, where each replica would generate a different token, so the file stays a template and the shipped-config test asserts it is refused for that reason. -
The admin API is no longer unauthenticated on loopback. It purges the cache, forces renders and drives warmup, and it ran open whenever it was bound to loopback — a boundary that does not exist inside a container, where the address is shared with every sidecar and anything that gets a shell. PRISM now generates a bearer token for the run when none is configured and logs it.
Set
admin.bearer_tokento keep it stable across restarts. Probes are unaffected:/health,/readyand/startupnever required a credential, so orchestrators and the image's own healthcheck need nothing.insecure_no_authis now the only way to run without authentication, and the startup refusal it used to override is gone — there is nothing left to refuse. -
A redirect keeps the cache policy the origin gave it. The redirect path retained
Locationand dropped everything else, so a302markedno-storereached the storage decision with nothing to decide on — it was kept for PRISM's configured TTL and then advertised aspublic. Login, locale, canonical and retired-product redirects are exactly the ones an origin marks that way. Security headers on a redirect survive now too. -
The origin's
Ageactually reaches the cache. The arithmetic for inheriting it shipped in this release; the header did not. Origin responses are filtered through a preserve list that did not mentionage, so nothing ever reached the code that subtracts it, and every entry still started at zero. The age arithmetic is also saturating now —Ageis a number an origin can get wrong, and the previous form panicked rather than clamping.Time spent rendering is still not counted:
created_atstarts when the entry is stored and the document arrived a render earlier. That is bounded byrender.timeout_secs. -
admin.peerswithout a configured token is refused at startup. A generated token cannot be shared, so each replica signed forwarded purges with its own and every peer rejected them — the fan-out silently did nothing and the other replicas served stale renders until TTL expiry, which is the failurepeersexists to prevent. Set the sameadmin.bearer_tokenon every replica. -
Liveness no longer mints a Chrome tab per call.
/healthis exempt from admin authentication so orchestrators need no credential, and the check it runs creates and closes a tab. A verdict is now reused for five seconds, which bounds a flood to one tab per window while staying well inside any probe interval. The cheap checks in front of it — crashed pool, no usable slots — are not cached and still answer immediately. -
The DNS guard fails closed. It exists because PRISM cannot see what Chrome's own resolver will return, and it passed through every hostname it could not resolve — on the reasoning that the request would fail anyway. That holds when a name genuinely does not resolve, and not when it resolves for Chrome but not here, from a different cache or a different server, which is the case worth guarding against.
An unresolved verdict is never cached, so the availability cost stays with the single request rather than taking the host out for the resolver TTL.
-
A forged
Hostno longer reaches the origin asX-Forwarded-Host. PRISM overwritesHostwith the origin's own name, so the public one is passed along separately — and it was passed along exactly as the client sent it. A storefront deriving its base URL from that header would build canonical links, redirects and sitemap entries pointing at a host the visitor chose.Where
public_base_urlis set, the operator has said which host this instance serves, and anything else is replaced with it. Where it is not, the client'sHoststill travels: PRISM is then the edge as far as it knows, and refusing it would break every single-hop deployment.
Documentation
-
The
allowed_originsexamples now match something. Every example on the security page was a URL, and the matcher compares bare hostnames — so each documented configuration matched nothing, and matching nothing blocks every subresource rather than failing loudly. An operator following the page got a rendered document with no CSS. A test now fails if an example regrows a scheme or an uppercase letter. -
The published claims were checked against the code. "Zero unsafe code" was false — one
unsafeblock exists, a libc page-size lookup — "631 tests" was stale by four hundred, and "3 independent security audits" contradicted this repository in writing, which says internal review rounds do not count as independent. All three are corrected and a test now fails the build if any returns. The footer's Privacy Policy and Terms links pointed at#; they now point at the published documents. -
rendering/status-codes.mdclaimed that not caching 5xx was the default whenerror_ttl_secsdefaults to 30, so anyone who wanted errors uncached and trusted the page got the opposite. The settings table inconfiguration/cache.mdis now parsed and compared againstCacheConfig::default(), so the reference an operator reads cannot drift from the binary again. -
ROADMAP.mdwas describing work that had already shipped in 1.5.3, and an external assessment scored cache correctness partly on that. It now records what is done, and sequences the route from the assessment's 7.5 to 9 as three releases — with the ceiling stated, since an independent penetration test, named customers and a second maintainer are not things code can deliver. -
Five fixes that had been applied to one surface and not its twin: the per-tab figure in the capacity model, the
#[ignore]d test count in README (173 → 223, measured), the "you may need to disableblock_private_cidrs" advice that survived in troubleshooting.md after SECURITY.md dropped it,WarmupStatus::Failedwhich the admin API reference never listed, and the supported-versions table still naming 1.4.x as current. -
The roadmap records Phase 1 as delivered, item by item with what makes each checkable, and names which of the six ranked deal-breakers it closes and which remain.
-
Three pages described behaviour this release changed. The README still said startup refuses an unauthenticated admin API on a non-loopback address; the admin page still said a null token leaves the API unauthenticated; and the cache page still said authenticated requests merely skip caching rather than bypassing rendering. All three now match the code, and the shared-token requirement for
admin.peersis written down where an operator configuring replicas will meet it. -
trusted_proxiesis documented as what it does. Three places said an empty list trusts all proxies "for backward compatibility". The code has always meant the opposite, and its tests pin that. Following the documentation was expensive: behind a CDN every visitor resolves to the same edge address, and the per-IP render limit — three by default — then applies to that one address on behalf of the whole site.
CI
- Wire-level Hurl contracts and rendered deployment manifests are now release prerequisites for the tag-triggered release path.
Internal
-
The 1.4.0 admin-auth default is pinned by a test that fails when it reverts. It lived only in untested main.rs wiring while the area's one test asserted a token-less server serves everything — reading as though open-by-default were the contract. The default path is now walked from
PrismConfig::default()through the same resolution main uses; the old test remains as what it is, the explicit opt-out contract. -
Production publication has teeth. An unscoped
|| truemade the promotion ssh chain unfailable (every rsync could fail, step green) — in the very step hardened the day before. Production verification now runs under pipefail and asserts the released version in the apt and apk indexes before staging is deleted; reachability alone read a five-months-stale apk index as healthy. Docker:latestmoves only after those assertions pass, not in the Package phase. A structural test pins all three against executable lines, not comments. -
A provenance guard that could never have run.
ci.ymlrefused to publish a tag not on main, and fetched withgit fetch --no-tags --depth=0— which git rejects outright as "depth 0 is not a positive number". The guard was added between releases, so its first execution was the 1.4.0 tag: it had never run once and failed where failing was most expensive.release.ymlcarries the same guard and fetched correctly, which is the only reason that tag was still publishable. A test now reads both copies. -
The apk index is rebuilt when promoting to production. The production APKINDEX was last written in March, so every .apk published since — 1.1.0 through 1.4.0 — was uploaded and stayed invisible:
apk add prismserved 1.0.0 while 1.4.0 sat beside it on disk. The promotion rsync runs with--ignore-existing, which is right for package files and wrong for an index, and the metadata script is invoked with--skip-apkfor production as well as staging, so nothing rebuilt it either.Staging was correct throughout, which is why release verification never caught it — that job installs from staging, the one repository whose index was current.
The production index is still stale until it is rebuilt on the host once; this fixes the next release, not the current state.
-
Publishing fails closed. The staging upload, the metadata rebuild and the repository verification each carried
continue-on-error: true, so a failure in any of them left the job green — and the GitHub Release ran onif: always()without depending on the repository verification at all. A release could therefore be announced against packages that were never delivered, which is what happened twice this week: the Docker image existed, the packages and the Release did not, and nothing failed.The release now waits for the repositories to be proven installable, and that proof asserts the installed version equals the tag rather than merely printing something. The package host's SSH key is pinned from a secret instead of being learned with
ssh-keyscanon every run. Two tests read the workflow and fail if either gate is removed.This needs a new repository secret,
PKG_DEPLOY_HOST_KEY, before the next release. Generate it once from a trusted network withssh-keyscan -H <package-host>. -
unknown-gitis enforcing. The allowlist named both git dependencies and then only warned about anything else, which is a signal nobody had to act on. -
CI runs the Rust gates when what compiles the Rust changes. The change filter matched
**/*.rs,Cargo.tomlandCargo.lock, so arust-toolchain.tomlordeny.tomledit skipped every one of them — which is exactly backwards, since those decide which lints run and which dependencies are allowed. -
read_body's error is boxed. Rust 1.98 widenedclippy::result_large_errto catch it: the error variant is a wholeResponse, 144 bytes carried by everyResultincluding the success path that never uses them. No behaviour change.Worth noting for whoever hits it next — CI pins its actions by SHA but resolves the Rust toolchain as
stable, so a release of Rust can fail the build with no change to this repository. That is what happened here.
1.3.3 — 2026-08-19
Fixed
-
The internal origin survived in JSON-LD. PRISM replaced it in attributes —
href,src,action,content,poster,data-src,srcset— and nowhere else. Structured data lives in the text of a script element, so none of that reached it, and a crawler was handed an entity identifying itself by a hostname that resolves nowhere:"url":"http://pwa-app:3000""target":"http://pwa-app:3000/search?q={search_term}"It is the part of the page a search engine trusts most and the one part that was not rewritten. Blocks are now parsed, their URL-valued strings rewritten at any depth, and re-serialised.
Two boundaries are deliberate and pinned down by tests: only strings that are the origin are touched, never one mentioned inside prose — editing inside a sentence would be rewriting the page's content — and only
application/ld+json, because a string in ordinary application code may be a cache key or a comparison whose meaning would change.Parsed rather than string-replaced because a blind swap can leave a block unparseable, and a search engine discards a block it cannot parse: a corrupted rewrite is worse than the wrong host it was fixing. Anything that does not parse going in is returned untouched.
Affects storefronts that derive JSON-LD URLs from
window.locationand setpublic_base_url. A Magento storefront building them from the store's configured base URL — most of them — was never affected.
1.3.2 — 2026-08-18
The other half of the warmup bug: 1.3.1 stopped entries from being refused, this stops them from being unreachable.
Fixed
-
Warmed cache entries could not be found by the traffic they were warmed for. Warmup rendered every URL and stored it,
/statuscounted the entries, and every crawler request missed and rendered the page again — filing a second entry beside the first. Measured on the reporting host: warmup stored 5 of 5 URLs with no errors, and a Googlebot request for one of them was a MISS.Acceptwas folded into the cache key. It is appended to every request so Chrome negotiates with the origin the way the client would, which is a rendering concern; the same list also fed the key. Live traffic always sendsAcceptand looked under<url>|vary:accept=html, while warmup and the admin/renderendpoint send no headers and filed their work under a bare<url>. Both sides worked and never met.The key now carries only headers listed in
[cache] vary. Nothing is lost: a client that refuses HTML is turned away before the cache is consulted and a passthrough is never stored, so every entry holds the same rendered HTML whateverAcceptasked for — the dimension split the cache three ways for identical content. An operator who listsacceptstill gets it, as a bounded class, and now gets it on requests that send noAccepttoo, which is what that class already meant.No configuration change is needed. If warmup appeared to do nothing on 1.3.1, this is why.
Internal
Accept-Encodingbehaviour is now covered by tests. It was already correct — one entry carries both the plain and gzip forms and the encoding is chosen per request, while the header is still forwarded to the origin on a passthrough — but nothing held it in place next to a change to how the cache key is built.
1.3.1 — 2026-08-18
A cache that filled at random, and an HTTP/2 advisory that landed the day after 1.3.0 shipped.
Fixed
-
A page's cache headers were taken from whatever it fetched, not from the page. Reported from the field as: warmup renders every URL,
/statusshows the entries, and live requests still miss. Measured against the affected origin, PRISM answered withno-store, no-cache, must-revalidate, max-age=0for a document whose origin returnss-maxage=60on forty consecutive requests — that header belonged to/graphqlon the same origin, which is what PHP sends oncesession_start()runs.PRISM read the document's status and headers from
wait_for_navigation_response(), which returns the main frame's most recently finished request. Any request the page makes overwrites that, with no filter for resource type, so a storefront calling its own API decided its own cacheability by coin toss. It now identifies the document by resource type fromNetwork.responseReceived.Any headless storefront whose pages call a session-bearing API was affected — Magento with a JS frontend most of all. If you saw a cache that never filled despite successful renders, this was it; no configuration change is needed.
Security
- RUSTSEC-2026-0258 —
h20.4.15 accepts unbounded empty DATA frames, so a peer can hold a stream open and force allocation without sending payload. Advisory published 2026-08-17, one day after 1.3.0. Updated to 0.4.16.
1.3.0 — 2026-08-17
Security and correctness fixes from an independent review of 1.2.0, plus the proxy-body streaming change that makes this a minor rather than a patch release. Several items change behaviour; read Upgrading before rolling out.
Upgrading
-
Varyno longer advertisesX-Prism-Variant; it advertisesUser-Agent.Varynames request headers, andX-Prism-Variantis one PRISM writes onto the response — so caches looked for a header nothing sends, found it absent every time, and collapsed bot/human and desktop/mobile into one object.User-Agentis what actually varies the body. It also caches poorly, which is the honest cost of being correct: to get useful edge caching, normalise the User-Agent at your edge into one header and list it in[cache] vary. That path was documented but never worked, because nothing told PRISM to vary on the header the edge was injecting. -
Origin redirects are now forwarded instead of followed. A 302 reached clients as
200 OKcarrying the target's body under the original URL. Bots now receive the redirect. Expect one extra round trip per redirected URL, and expect previously-duplicated URLs to consolidate in search results. -
Chrome no longer accepts invalid TLS certificates. If your origin serves HTTPS with a self-signed or expired certificate, renders will now fail. Trust the CA, or set
render.accept_invalid_certs = truedeliberately. -
block_private_cidrs = truein the shipped Docker and Kubernetes configs. The configured origin was always exempt from this check, so it never needed to be off; with it off, a rendered page could reach any other private address, including PRISM's own unauthenticated admin API. -
prism validatenow rejects configurations it used to accept: aserver.originthat is not an absolute http(s) URL, andwarmup.concurrency = 0. Both started successfully before and failed later — the second by panicking the process on the first warmup.
Fixed
-
The image's health check could wedge its host. It probed
/healthevery ten seconds./healthdoes a real CDP round-trip, so it created and closed a browser page six times a minute in competition with actual renders — and once the browser was loaded enough that the probe outran its timeout, the daemon killed it. On current Docker and containerd a killed exec strands its state on the/runtmpfs: an<exec-id>.pidfile plus-stdout/-stderrFIFOs, none of which are ever reaped. That made the probe self-reinforcing, and one host accumulated 198k pid files — a 4 KB tmpfs page each, so the entire 776 MB of/run— with 396k FIFOs beside them holding 41% of the inode table. containerd could then not write init state and no container on the machine could start, including ones unrelated to PRISM. It presents asno space left on deviceon a host whose disk has plenty of room, because the full filesystem is a tmpfs. Now probes/ready, which reads atomics and cannot hang, every 30 seconds with a 60-second start period. A probe that completes normally leaks nothing, so this closes the leak rather than slowing it; the deployment docs carry the host-side cleanup for hosts that already have a backlog. -
Accept-Encoding: gzip;q=0was answered with a gzip body. The header was matched by substring, so an explicit refusal read as acceptance.notgzipmatched,GZIPdid not, and*did not. -
Cache byte accounting was updated outside the lock that decides eviction. Two writers racing on one key could subtract an entry's bytes before they were added, wrapping the unsigned total to ~
usize::MAX— after which eviction emptied the cache on every store, silently in release builds. -
Purges reached one pod behind a headless Kubernetes Service instead of all of them. Each peer is now resolved and purged at every address, so one Service name covers an autoscaled fleet.
-
The internal origin leaked into rendered output. Anything a page derived from
window.location— a JS-generated canonical, anog:url— named the internal address and was served to crawlers as the site's identity. -
The Kubernetes manifest referenced a registry this project does not publish to with a digest of all zeroes, so applying it produced
ImagePullBackOff. -
The origin-timeout log reported "30s" regardless of
proxy_timeout_secs.
Changed
-
Proxy bodies stream instead of buffering. The pass-through path held each request and response body in memory so headers could be rewritten — headers that live in
partsand never needed the body at all. At the shipped ceilings that was up to 10 MiB of request and 50 MiB of response per connection against a 10,000-connection limit, with nothing bounding the aggregate.Every limit that mattered is still enforced.
Content-Lengthis refused up front in both directions, andLimitedcaps a body that under-declares its length or omits it entirely.Behaviour change:
server.body_read_timeout_secsbecomes an idle timeout — the longest a body may stall between frames — rather than a total deadline on reading it. That is the bound streaming wants: a total deadline cannot tell a slow attacker from a large legitimate upload, so raising it restores slowloris and lowering it severs a real upload over a poor link. A size violation now also surfaces mid-stream rather than before any bytes are sent; the statuses are unchanged (413 too large, 408 stalled).This is why the release is 1.3.0 and not a patch: it rewrites the request hot path and changes documented behaviour, and neither belongs in a patch version that operators would reasonably take without reading.
-
render.auto_scroll(defaulttrue) turns off the lazy-load scroll pass. On a tall page it spends its full five-second budget on every render, which sites that hydrate up front were paying for nothing. -
render.accept_invalid_certs(defaultfalse) is the TLS escape hatch. -
Test builds verify licences against a separate throwaway key (
test-license-keyfeature), so no production-valid credential exists in the repository. -
The release is now gated on the
corpus,scale-smokeandcoveragejobs, which describe themselves as required but were not inneeds.crossis pinned by commit rather than tracking its default branch. CI cancels superseded pull-request runs.
1.2.0 — 2026-08-10
Upgrading
Four changes alter behaviour on upgrade even though nothing breaks at the API level. Read these before rolling out.
-
Memory: 4 GiB minimum for the default eight-tab pool, up from a documented 2 GiB. Four warm tabs measure ~1.5 GiB before any large render, so the old figure was below the working set of half the default pool. Anyone running to the previous recommendation is under-provisioned and will discover it under load. Lower
render.pool.tabsbefore lowering the limit. -
Vary: Acceptis now emitted. PRISM varies its own cache on the representation class, so a CDN must be told. This changes cache fragmentation for every deployment behind one — expect a lower hit rate until the edge re-warms. -
Non-HTML routes proxy the origin instead of returning 503. Correct — there was nothing to render — but an uncached feed now reaches the origin about three times per bot request where it previously produced a cheap error. Watch
prism_passthrough_not_html_totalafter upgrading. -
render.max_dom_nodesdefaults to 400,000. A new rejection path that did not exist before. Documents above it are refused before serialisation.
Known limitations
- A non-HTML passthrough costs roughly three origin requests per client request.
render.timeout_secsabove 30 requires the patched chromiumoxide pinned inCargo.toml. Building against the published crate silently reimposes a 30-second ceiling on every CDP command;tests/cdp_timeout_e2e.rsfails if that happens.- Chrome intermittently stalls a single navigation and is retried, not
fixed. On roughly one CI run in four, one route's
Page.navigatereceives no answer — no response, no events, no CDP traffic at all — while the browser stays healthy and the next request renders the same fixture in about two seconds. It affects a different local fixture each time, with no crash, recycle or circuit event in the run. The cause is not known. PRISM now discards the tab and retries once, which recovers the request; watchprism_render_retries_total, because a climbing counter means the stall is spreading and the mitigation is hiding it. - Not yet evidenced: commerce-scale validation, a Rendertron comparison, and a sustained soak. Suitable for a controlled rollout on that basis.
Added
-
cache.honor_request_no_cache(defaultfalse). When enabled, a request carryingCache-Control: no-cache,Cache-Control: max-age=0orPragma: no-cacheis re-rendered instead of being served a stored response, per RFC 9111 §5.2.1.4, and the fresh render is stored so it refreshes the entry every later request sees.It is off by default on purpose: PRISM is a shared cache usually reachable from the public internet, and a revalidation costs a full Chrome render, so honouring the header unconditionally is a cache-busting denial of service that needs no credentials. Enable it only where callers are trusted, or behind a CDN that strips the header from client requests.
Added
prism_render_retries_total— renders retried after a stalled navigation. One request costing two tabs, counted separately fromprism_render_totalso the mitigation above cannot hide its own cost.
Fixed
-
A stalled navigation is retried once on a fresh tab. Chrome occasionally stops answering on one tab while remaining healthy; the tab was already being discarded, so a second attempt gets a clean one and generally succeeds. Only a timeout is retried — a systemic fault would get twice the load, a capacity refusal would steal a queue slot, and a page-local verdict would re-derive the same answer. See Known limitations: this recovers from the stall, it does not explain it.
-
A client that refused HTML was served a warmed cache entry. The
Acceptparser ranked quality weight above specificity, sotext/html;q=0, */*;q=1— a client asking for anything but HTML — was read as accepting it atq=1. RFC 9110 §12.5.1 is the other way round: the most specific matching range decides, and the broader ones speak only for the types the client did not name.The cache is what made it serious. Such a request was classed
html, the same class an ordinary browser warms, so on a warm entry it was handed a stored render with no origin contact and no render at all — nothing on the render path was in a position to notice, and a cold-cache test passes against it.Consulting
application/xhtml+xmlwas the other half: PRISM answersContent-Type: text/htmland nothing else, so asking whether a client accepts a type PRISM never sends is how*/*came to override a refusal that namedtext/htmloutright. Both the acceptability check and the cache class now ask about the one type PRISM emits, so they cannot disagree.Behaviour change: an
Acceptnaming onlyapplication/xhtml+xmlis now passed through rather than rendered. Browsers listtext/htmlalongside it, so no real client is affected. -
render.timeout_secsabove 30 did nothing. chromiumoxide 0.9.1 takes both the per-command future's delay and the frame manager's navigation deadline from a crate constant rather than the configuredrequest_timeout, so any page slower than 30 seconds failed as an opaqueNavigation("Request timed out")regardless of configuration — and the bundled test stacks asking for 60 and 180 seconds were describing a bound that could not exist.PRISM now builds against a fork of v0.9.1 carrying one commit that threads the configured value through to both sites, pinned by revision in
Cargo.toml. The fork point's tree was verified byte-identical to the published crate.This also explains a run of intermittent E2E failures previously written off as CI flakiness: always one route, always just past 30 seconds, always green on rerun, because the page usually finished inside the ceiling.
-
The runtime image installed a floating Chromium. Both base images are digest-pinned, but
apt-get install chromiumresolved against a live index, so rebuilding the same commit could install a different browser — the one component whose behaviour PRISM's serialisation, timing and size preflight depend on most directly. Now pinned by version via a build argument. -
Acceptwas matched by substring, not parsed.text/html;q=0, application/json;q=1contains the text "text/html" while explicitly refusing it, so PRISM read that as acceptance, rewrote the header it forwarded totext/html, and returned the rendered page the client had just rejected.Accept: text/plainwas likewise upgraded to*/*.There is now a real parse: media ranges with quality weights,
q=0treated as the refusal RFC 9110 defines rather than a low preference, the most specific matching range deciding the weight, and media-type parameters not mistaken for weights. Two decisions that were one are now separate — the client's own header is forwarded to the origin unchanged so negotiation stays theirs, while the cache is keyed on a bounded three-value class so the key cannot be inflated. A request that will not accept HTML is refused before a tab is taken. -
A background refresh of non-HTML content counted two passthroughs. Both
do_renderand the background arm recorded it, for a refresh that served no client at all. Only the foreground path counts now. -
Successful passthroughs were filed as size rejections. They incremented
prism_oversize_rejections_total{reason="not_html"}, reporting a ceiling problem where none existed. That label is gone; the passthrough counter records the outcome once, where a client is actually served. -
Forwarding
Acceptto Chrome made it a cache-busting dimension. The previous release forwarded the client's raw header so Chrome and the origin fallback would negotiate the same way. It was also keyed on verbatim, and the header is client-controlled and effectively unlimited —text/html;q=0.9,*/*; q=0.8,x=<nonce>is a valid HTML-compatible request — so anyone could force a fresh Chrome render per request and evict the cache on the way.The cache key is now a bounded three-value representation class rather than the raw header, keeping cardinality fixed while still distinguishing negotiations that produce genuinely different documents. The header sent to the origin is the client's own, unchanged — see the parsing entry above, which supersedes an earlier version of this fix that forwarded the class and so asked origins for representations clients had not requested.
Acceptis also declared in the responseVary. PRISM varies its cache on it, so a CDN not told about it would collapse those variants and serve HTML to a JSON client — the same defect already fixed once for configured dimensions. -
Non-HTML passthroughs were still counted as errors in two places. The authenticated-skip branch recorded
requests_error, and background refresh recorded a page-local render failure. Both now go through oneis_client_successpredicate rather than an inline check per call site, which is what let the miscount reappear twice after being fixed once. -
A non-HTML passthrough was counted as an error and could serve the wrong representation. Refusing to render a non-HTML document is correct, but the request that follows is a success: the client receives the origin's own response with its own status. It was being recorded as
requests_errorand as a page-local render failure, so a site with one busy feed read as broken on every dashboard. Passthroughs now have their own neutral counter,prism_passthrough_not_html_total.The two origin requests a passthrough makes — Chrome's navigation, then the fallback — were also not equivalent. Chrome received only the configured cache dimensions, so a content-negotiated endpoint could hand it JSON while the fallback, which forwards the client's real headers, was given HTML and served it unrendered.
Acceptis now always forwarded to Chrome, whether or not it is a configured dimension, so both requests see the same representation.Known cost, measured rather than assumed: one passthrough currently makes three origin requests.
prism_passthrough_not_html_totalexists so that volume is visible, and the test asserts the count stays small and bounded. -
Documents are now bounded before they are fully materialised, and the bound cannot be walked around.
render.max_html_byteswas enforced only after Chrome had built the whole document, serialised it, pushed it across CDP and Rust had allocated its own copy — the peak the limit exists to prevent, paid three times before anything checked it.A preflight now walks the DOM inside the page and adds up what serialising each node would cost, stopping the moment the budget is exceeded and never concatenating: a string longer than the remaining budget is over by inspection and is never read, so scanning stays bounded even on a hostile page. Alongside it,
render.max_dom_nodes(default 400,000) caps element count.Getting that estimate right took several passes, each closing a way past it:
- attributes, comments, tag names and escaping expansion are all priced, not just visible text — a page whose weight is one 40 MiB attribute scored as almost nothing under the first version;
- UTF-8 is counted exactly, including surrogate pairs, and unpaired surrogates as U+FFFD, so an emoji is four bytes and a lone surrogate three;
- element names are counted in UTF-8 too, since Chrome accepts non-ASCII names and ten million two-byte characters is 20 MB in each of the opening and closing tags;
<template>contents are walked as additional roots — they live in a separateDocumentFragmentthat no walker rooted atdocumentElemententers, whileouterHTMLemits every byte of them;- processing instructions are priced, and CDATA sections as the escaped text
the HTML serialiser actually emits — verified against a live render, which
produced
&per character and no<![CDATA[wrapper; - a node type with no rule of its own refuses the page rather than being guessed at, because the traversal's whole value is knowing what the serialiser will write, and for an unanticipated type it does not;
- documents that are not
text/htmlare not rendered at all. Every rule here describes the HTML serialiser; XML and XHTML serialise differently, and the same traversal would mis-price them in the under-counting direction. Such a response is proxied from the origin unchanged rather than refused — there was nothing to render, and the origin already produced the right answer. - comments and raw-text elements (
<script>,<style>) are charged verbatim rather than with escaping expansion, which had inflated them enough to refuse documents comfortably inside the budget; - a preflight that cannot complete — an evaluation error, an unreadable result, a timeout — refuses the page instead of falling through to serialising it, since the traversal runs inside the page and a hostile document can reach that path deliberately.
The exact byte check after extraction remains authoritative; the preflight exists to avoid reaching it.
prism_oversize_rejections_total{reason}records which check refused each document, so it is visible if the cheap ones stop earning their place. -
Three
nanoidadvisories were fixable and had not been fixed. 3.3.16 intests/e2e/spa-fixture,about/websiteanddemo/magento-pwa/pwa-app, all upgraded to 3.3.18. The accepted-advisories note previously claimed every open alert resolved toimage-size; that was written after auditing only two of the projects and is corrected. -
Shipped memory guidance disagreed with itself and with the default pool. The README recommended 2 GiB in three places while the packaged systemd unit set 4G and Kubernetes requested 3Gi — for a default of eight tabs, where four warm tabs alone measure about 1.5 GiB before any large render arrives. A limit below the working set is an OOM kill during the first spike, and the figure an operator copies is whichever file they opened. Everything now states 4 GiB, with the measurement basis and how to budget from real traffic, and a test fails if the surfaces drift apart again.
-
The Chrome memory gauge could report a process that no longer existed.
prism_chrome_rss_byteswas written once per completed render, just before a possible recycle, and never afterwards — so following a restart it kept the outgoing browser's peak while the new one was a fraction of the size, and a shut-down pool kept reporting its last live reading. A gauge is read as "now"; this one could be minutes old with nothing to say so.It is now sampled every five seconds by a supervised task, refreshed immediately after a restart, and zeroed on shutdown. The new
prism_chrome_rss_age_secondsmakes staleness visible instead of implicit, and the HELP text states that the value covers the browser process and all its descendants. -
A page that could not render opened the circuit breaker for every other page. Each render failure was fed to the shared Chrome circuit regardless of what it meant, so a document rejected on purpose by
render.max_html_bytes— requested repeatedly, which is what a crawler does — tripped the breaker and stopped rendering site-wide while Chrome was in perfect health.Failures are now classified:
systemic(the browser or its CDP connection is broken),capacity(no slot, or maintenance in progress),page_local(this URL cannot render) andpolicy(refused by configuration). Onlysystemicfeeds the circuit. The rest are counted in the newprism_render_failures_total{class}so being circuit-neutral does not make them invisible, and a page-local failure returns the half-open probe instead of stranding it, which would stall recovery.Verified against a live stack: twelve consecutive 33 MiB rejections against a failure threshold of five leave the circuit closed, all twelve counted as
page_local, and a healthy URL still renders. Killing Chrome outright is still detected and still recovers. -
Requests waiting out a browser recycle were not counted against
queue_max. Waiting rather than refusing is correct — a recycle is planned maintenance and the replacement is seconds away — but a waiter is still a queued request: it holds a client connection and will take a tab as soon as one exists. Queue depth was accounted per attempt, so a parked waiter left the count and the queue reported itself empty while an unbounded number of them accumulated.queue_max, the only bound between a slow recycle and running out of memory, applied to none of them. Queue admission now spans the whole acquire, waits included. -
A planned browser recycle was counted and recovered from as a crash, and returned 503 while it ran. Every exit reached the monitor through one untyped channel, so closing Chrome deliberately — which is what a recycle is — was indistinguishable from it dying. The recovery restart then closed a browser again, emitting another exit, which was counted again: five recycles produced five crashes and five restarts, with 503s throughout, and the crash counter said the browser was unstable when nothing had gone wrong.
Exits now carry the generation they belong to and why they happened. An exit is a crash only when it is the current generation, unexpected, and the pool is still running — a planned close is marked before it is made, and an exit from a generation that is already gone cannot restart the browser that replaced it.
Requests arriving mid-recycle now wait for the new browser instead of being refused. A recycle is planned maintenance and the replacement is seconds away; refusing during that window turned routine maintenance into an outage on every eighth render at the shipped
max_renders_per_browser. The wait is bounded by the deadline a queued request already has, so a recycle that never finishes still fails rather than hanging.Measured before: 5 recycles, 5 crashes, 5 restarts, 503s. After: 20 renders across 2 recycles, 2 restarts, zero crashes, zero 503s, circuit closed.
-
A request body that arrived too slowly was answered
413 Payload Too Large. Slow and oversized are different client faults; the read timeout now answers408 Request Timeout. Telling a slow client to send less does not help it, and conflating the two hid slow-client attacks among genuine size rejections in the logs. -
The admin listener accepted connections without bound. It serves liveness, readiness, metrics and purge, so a caller holding connections open could starve the health signal an orchestrator depends on. Capped at 64 concurrent connections; past that, connections are refused rather than queued, because an admin request that cannot be served promptly is worse than one that fails fast — the caller is usually a probe with its own deadline.
-
POST /purge/patternleft device and vary variants behind. The glob was matched against the whole cache key, which ends in|mobileor|vary:..., so a pattern naming a page — the shape an operator writes to purge one URL — matched only the desktop, no-vary entry. The response reported success while every other representation kept serving the content that had just been invalidated./purge/urlalready removed all variants; the pattern endpoint now does the same, and patterns written against a suffix still work. -
Cache dimensions PRISM keys on were missing from the response
Vary. With[cache] varyconfigured, PRISM fragmented its own cache correctly but advertised onlyX-Prism-VariantandAccept-Encoding. A CDN in front of PRISM was therefore free to collapse those variants and serve one language's render to another language's request. The configured dimensions are now merged intoVaryalongside the origin's and PRISM's own.
Testing
-
The scale suite asserted the wrong recycle metric and is now required in CI.
prism_tab_recycles_totalcounts per-render page and context replacement, which happens on every release regardless of browser lifecycle, so the recycle test passed on ordinary traffic. It now asserts onprism_chrome_recycles_total, requires restarts to equal recycles (more means a planned close was recovered from as a crash), requires the crash counter not to move, requires the circuit to stay closed, and requires the memory sample to be fresh after the restart.settle()fails instead of continuing when renders never drain. The smoke tier runs as a required job; the full matrix stays on its own schedule. -
The shared route corpus is now executed, and checked against the fixture.
tests/corpus/routes.tomldeclared in its own header that four suites consumed it while nothing in the repository referenced it. It had drifted accordingly:/products/localizedand/consent-variantwere not fixture routes at all (/localizedand/consentare), the soft-404 entry pointed at a path that only reached the catch-all, seven pages were marked "not yet implemented" long after they were built, and two of the product prices it asserted did not match what the fixture renders — never noticed, because nothing ever compared them.tests/corpus_manifest.rsnow checks the manifest against the fixture's router and data on everycargo test, in both directions, with no server required.tests/corpus_e2e.rsdrives every declared route through a real PRISM and checks the declared expectations. Both run in a required CI job. -
Device-aware rendering is exercised by the e2e stack.
[render.viewport]was absent from the e2e config, so every request rendered asbot-desktopand the fixture's mobile page served its desktop branch to a phone. The old assertion accepted the substringView, which both branches satisfy, so the suite passed with the feature switched off.
1.1.5
Supersedes 1.1.4, which was tagged but never published: the release pipeline checks out the tag it is releasing, so a tag cut before a packaging fix cannot carry it. This is the first tag containing both the rpm signing fix and the CI change that stops Chrome's container limitations failing a package check.
Packaging and documentation fixes. The Rust crate is unchanged apart from the version, but two of the three ways of installing PRISM did not work.
Fixed
-
dnf install prismfailed on every RHEL, Rocky, Fedora and CentOS host. The repository publisher signedrepomd.xmlbut never the packages.repo_gpgcheckvalidates the index;gpgcheck— which the shipped repository config sets — validates each package, and an unsigned rpm fails it with "Signature verification failed" however well signed the metadata is. Packages are now signed withrpm --addsign. -
apk add prismfailed on Alpine. The repository line was written as<url>/apk/<arch>, and apk appends the architecture itself, so it fetched.../apk/x86_64/x86_64/APKINDEX.tar.gzand reported "no such package" even though the index was published and reachable. -
The documented nginx configuration did not strip the bypass header it said it stripped.
proxy_set_headerdirectives are not merged across levels: alocationdefining even one of its own discards every one from the enclosingserverblock. The example placed the strip atserverlevel andproxy_set_header Connection ""insidelocation /, so a client could forgeX-Prism-Bypass, andHost,X-Real-IP,X-Forwarded-ForandX-Forwarded-Protowere all dropped as well.
Added
-
dist/reverse-proxy/ships complete nginx and Apache configurations, so they can be copied rather than transcribed out of a documentation page — which is how theproxy_set_headerplacement bug survived review. -
dist/integration-testing/run-distro-matrix.shruns the release's package verification locally against real packages. The matrix could previously only be exercised by cutting a release.
CI
- Package verification covers CentOS Stream 9, openSUSE Tumbleweed and Ubuntu 20.04, installs the arm64 packages under emulation, and runs the static tarball on alpine, debian-slim and busybox.
- Those tests now start PRISM — proxy, admin endpoint, a rendered bot request and a clean shutdown — rather than checking that files were laid down.
- The repository verification jobs install with
gpgcheck=1instead of0, so they exercise what a user's machine does. With it disabled they passed throughout the outage above. - PRISM builds and passes its browser-free tests on FreeBSD 14.
1.1.3
A review pass over the request path, plus the security work that came out of enabling code scanning. Most fixes here are paths that worked normally and failed on an abnormal exit — a panic, a cancelled future, a shutdown — or a check that recognised one spelling of a thing rather than the thing. No behaviour, configuration or interface changes; a 1.0.0 config still loads.
Security
-
Alternative encodings of internal addresses were not blocked. The private-address check understood
::ffff:127.0.0.1but not::127.0.0.1. Both name the loopback interface; the check usedto_ipv4_mapped, which by definition ignores the IPv4-compatible form. The same gap applied to 6to4 (2002:7f00:1::) and the well-known NAT64 prefix (64:ff9b::7f00:1), which wrap an IPv4 address inside an IPv6 one — including169.254.169.254.This runs on DNS results as well as URL literals, so a hostname resolving to any of these spellings previously passed both the rebinding defence and subresource filtering. The embedded address is now decoded and checked, so
64:ff9b::808:808still reaches 8.8.8.8 rather than the prefix being refused outright. Missing on the IPv4 side as well: broadcast, multicast, the 240/4 reserved block, and 0.0.0.0/8 — only the single address0.0.0.0was caught, not the range. -
A self-referential origin was only caught in four spellings. The check that stops PRISM rendering itself compared the origin host against
localhost,127.0.0.1,::1and0.0.0.0as text. Binding to0.0.0.0makes the whole 127/8 range reach the process, sohttp://127.0.0.2:4000against a0.0.0.0:4000listener was an infinite render loop that validation accepted. The host is now parsed as an address. -
The benchmark harness returned exception text — including stack and local paths — to clients on a failed render.
Fixed
-
A panic in a background re-render took that URL out of service. The foreground path has used an RAII guard since the coalescing rewrite; the background path cleaned up by calling
remove_inflighton each of its six exit branches, and a panic skips all six. The leak does not heal: the URL then looks permanently in flight, so it is never re-rendered, and once past grace every miss waits outtimeout_secs + 5and returns 503 — until the process restarts. -
Two clicks on
/warmupstarted two warmups. The running check and the state reset were separate critical sections, so concurrent callers all observed a non-running state and all proceeded. Eight concurrent callers all won, crawling the same sitemap and interleaving their counters through one shared state. A panic in the sitemap fetch or parse additionally stranded the status onRunningfor the life of the process, refusing every later warmup while reporting one in progress. -
A client hanging up mid-cleanup leaked a tab and its context.
releasemoved the page and context out of theTabGuardbefore its first await, making that guard'sDropa no-op for the rest of the function. Chrome kept both until the next browser recycle. -
Shutdown stopped accepting connections but not requests. The drain broke the accept loop and waited on the in-flight counter. An established keep-alive connection is just an open socket, and HTTP/1.1 lets the client send another request — which the server answered, for the whole drain window. Connections now shut down gracefully. Readiness already failed before the drain, so the hole was requests arriving on connections already handed over.
Performance
-
Post-processing parsed the document once per transform. Each enabled transform ran as its own
rewrite_str, so a page with everything enabled was parsed six times. Measured on a 1.9 MB document: 32.4 ms before, 16.7 ms after, on every cache miss. Not the 5.6x the parse count suggests — the handlers cost real time themselves. -
Content validation copied the document to look at it.
extract_titleallocated a full lowercased duplicate to find one tag;strip_html_tagsallocated aStringfor every<; the visible-text check built aVecof every word and joined it into a newStringto use only its length. Measured on a 1.46 MB page: 3.33 ms before, 1.80 ms after, per render. -
A rewrite failure in post-processing was silent: the input was returned unchanged, so a page whose transforms did not run was served to bots with the scripts this exists to strip. It now warns.
Testing
-
IPv6 is covered end to end rather than only as a predicate over
IpAddr: the address cases run through URL parsing, which rewrites these hosts ([::127.0.0.1]normalises to[::7f00:1]), and a new test drives the real binary bound to[::1]. Nothing previously bound an IPv6 socket anywhere. -
The keep-alive shutdown test was rewritten after it raced on CI. Two further attempts passed with the fix reverted — once because the process exited before the assertion, once because the request under test hit a stalled origin and "no reply" reads like "refused". It now runs an origin that stalls one path and answers another.
Security tooling
-
CodeQL analyses Rust, the JavaScript projects and the Actions workflows. Secret scanning runs with push protection and generic-pattern detection.
mainrequires a pull request with passing checks;v*tags cannot be moved or deleted. Every workflow declarespermissions:explicitly and the release phases run with a read-only token. -
Eight CodeQL findings fixed in the demo and e2e fixtures: incomplete single-pass HTML stripping replaced with a scan, and ten committed build artifacts removed.
1.1.2
Fixed
-
CDP requests now time out on PRISM's schedule, not the library's. chromiumoxide defaults every CDP request to a hard-coded 30s that owes nothing to
render.timeout_secs. Any configuration above ~10s therefore had a ceiling it did not ask for: a slow navigation failed as an opaque "Request timed out" at 30s, before PRISM's own global budget applied, and short-circuiting the handling that closes and discards a possibly-wedged tab. It surfaced as intermittent render failures under load.The CDP request timeout is now derived from the global render budget, so PRISM's own bounds always decide first. It is never sized below chromiumoxide's default, so short configurations are unaffected.
Operators who raised
render.timeout_secsabove 30 were silently not getting it. They are now.
Changed
- The demo app moved off
react-router-domtoreact-routerv8. The former has no v8 and pinsreact-router@7, which made the patched version unreachable. Not shipped in PRISM; listed for completeness.
1.1.1
Dependency and supply-chain maintenance. No changes to PRISM's behaviour, configuration, or public interfaces — the Rust crate is untouched apart from the version bump.
Security
Dependabot alerts and automated security updates were enabled on the
repository, which surfaced 77 advisories (2 critical, 32 high). None were in
Rust — cargo audit reports zero vulnerabilities across the dependency
tree, and the shipped binary and container image were never affected. All were
in four JavaScript sub-projects that ship with the repo but not with PRISM:
the docs site, the marketing site, the e2e render fixture, and the demo app.
Open advisories reduced from 77 to a small remainder, both criticals cleared
(websocket-driver, shell-quote). Fixes came via Dependabot PRs, npm
overrides for transitive packages it could not reach, and two migrations:
the demo to react-router-dom 7, and the e2e fixture to react-router v8.
CI
- Build jobs for the four JavaScript projects. None had any CI, so every dependency bump against them was unverifiable — a broken build would only have surfaced on deploy. Two of these sites are publicly deployed.
- E2E Docker jobs now run when their own inputs change. They declared
needs: build, andbuildonly runs for Rust changes, so a skipped dependency silently skipped all three jobs — meaning changes undertests/e2e/**or to a Dockerfile never ran the tests that exist to cover them. They build their own image and never used that artifact. - Dependabot extended to the four npm projects; it previously watched only the repository root for cargo, actions and docker.
tests/shipped_configs.rsloads every config file the repo ships, so a validation change that invalidates one fails in seconds locally instead of after a full multi-arch release build.
Fixed
Several upgrades are only valid as a pair, and Dependabot proposes each half separately. Merging them individually broke all three JavaScript apps:
about/website: react 19 landed without next 16 (next 14 peers on react ^18). Completing the pair then required moving the Google Fonts@importabove the@tailwinddirectives — Next 16's Turbopack enforces the CSS spec where webpack did not — and pinning TypeScript to 6, since 7 is the native rewrite and no longer exposes the compiler API Next type-checks with.demo: react-dom 19 without react 19, and@vitejs/plugin-react6 (peers on vite ^8) without vite 8. Apollo Client 4 additionally moved its React bindings out of the root entry, so seven files needed repointing at@apollo/client/react.devdocs: the Docusaurus 3.10 bump requires@docusaurus/fasteras an explicit dependency whenfuture.v4is enabled; 3.9 bundled it.tests/e2e/spa-fixture: react 19 without react-dom 19.
1.1.0
Config schema is backward compatible — a 1.0.0 config loads unchanged, since no keys were removed and every new key has a default. Several defaults changed behaviour, though, so read "Changed defaults" before upgrading.
Changed defaults
| Setting | 1.0.0 | 1.1.0 | Effect |
|---|---|---|---|
security.rate_limit_per_ip | 0 (off) | 3 | Renders above 3/s from one IP get 429. Charged on cache misses only, so cached traffic — including everything behind a CDN sharing one source IP — is unaffected. Set to 0 to restore the old behaviour. |
render.pool.max_browser_age_secs | — (no recycling) | 3600 | Chrome is recycled hourly. Each recycle refuses renders for ~370 ms. |
render.pool.max_renders_per_browser | — (no recycling) | 5000 | As above, whichever comes first. |
render.circuit_breaker | consecutive failures only | + failure-rate window | The breaker can now open on sustained partial failure. Previously a renderer failing 75% of requests never tripped it. |
| Origin allowance | host + port | scheme + host + port | origin = "http://x:8080" no longer also permits https://x:8080. |
systemd MemoryMax | 2G | 4G | Sized for the shipped render.pool.tabs = 8. PRISM now warns at startup if the pool cannot fit the detected memory limit. |
Fixed
- Chrome crash recovery never fired. The only crash signal was the CDP handler stream ending, which never happens while the pool holds the browser handle — so a dead Chrome was never detected and the pool never recovered. Replaced with an active liveness watchdog; detection now ~13 s. It also catches a browser that is running but no longer servicing CDP.
- Cache-key collision across variants. Vary values were joined without escaping, so a client could craft a header value that collided with a different variant's key and be served that variant's HTML. Components are now escaped.
- Transient
accept()errors killed the process.EMFILEunder load, or a client aborting during the TCP handshake, terminated the accept loop and exited. Now classified: retry, back off, or stop only for an unusable listener. - Panic on an origin-controlled header. A
Varyresponse header containing a control byte panicked the connection task. utm_*did not strip every matching parameter. The matcher used shell glob semantics where*stops at/, soutm_a/bsurvived into the cache key.- Circuit breaker: a single success no longer force-closes an open circuit.
HEADerror responses no longer carry a body (RFC 9110).- Page-supplied status codes (
status_from_meta) are clamped to 200–599. purge_patternno longer holds the cache lock across its whole scan.
Added
- Cross-replica purge.
admin.peersfansPOST /purge/*out to sibling replicas, so one purge invalidates the fleet. Forwarded purges are never re-forwarded, so peers may list each other. - Readiness and startup probes.
/readyand/startupanswer from atomics (safe to poll often);/healthkeeps doing a real CDP round-trip and is the liveness signal. Readiness fails at the start of the drain so a load balancer can route away before requests are cut off. - Correlation IDs.
x-request-idis accepted from upstream (validated) or generated, recorded on every tracing span, and returned on every response. - Subresource DNS-rebinding defence. Resolution is now validated for every subresource a page requests, not just the navigation URL.
- Metrics: queue-wait histogram, rejections by reason, render phase timings, Chrome process-tree RSS, background-task failures, proactive recycles.
- Background tasks are supervised — one dying is now logged and counted rather than silently removing the capability it provided.
- Chrome restart backoff is jittered, so replicas failing together do not retry in lockstep.
- Kubernetes manifests (
deploy/k8s/) with probes, PDB, HPA on queue depth, and an egress NetworkPolicy. - Benchmark suite (
bench/): self-contained comparison against Rendertron and a Puppeteer baseline.
Performance
- Cache hits no longer re-compress. Compression happens once at store time; a hit is a refcount bump. Measured 3.07 ms → effectively zero for a 1.1 MB page.
- Route matching 1896 ns → 48 ns per request by classifying patterns once instead of running every glob.
- URL security checks parse the URL once instead of twice (335 → 182 ns).
- Tab release no longer serializes on an exclusive browser lock.
Dependencies
chromiumoxide0.7 → 0.9, which also removed a duplicatedreqweststacklol_html2.7 → 3.0,rand0.8 → 0.10,lru0.16 → 0.18,toml0.8 → 1.1- Total dependency count 317 → 298
- Base images pinned by digest; builds use
--locked
1.0.0
Initial release.