Skip to main content

Changelog

Every release, and for the ones that change behaviour, what to do about it. Look for the Upgrading section of the version you are moving to — that is the migration guide, and it names the things that are visible from outside PRISM.

1.4.0 — 2026-09-03​

The version line was renumbered here. Everything below shipped previously as 1.4.0 through 1.7.1 across eight tags in ten days — a pace that said more about how the work was sequenced than about how much of it there was. Those numbers are withdrawn from the package repositories and the registry; this is the release that replaces them, and it contains all of their substance.

Nothing is lost: each entry is the one written when the change was made. They are grouped by kind rather than by the release they happened to land in.

The git tags for the withdrawn versions still exist and cannot be removed — the tag ruleset refuses deletion. They are source history, not a distribution channel; installs come from apt and the registry.

Upgrading​

No configuration change is required. Three corrections can make previously hidden origin behaviour visible:

  • Repeated response fields are preserved as repeated fields. An origin sending two Cache-Control, X-Robots-Tag, Link, Vary or security-header fields now has both honoured instead of a newline-joined value being misparsed or dropped.
  • Parameters listed in cache.strip_query_params no longer reach Chrome. They were already absent from the shared key, so allowing them to shape the stored document let the first tracking URL decide what every later crawler saw.
  • Sitemap indexes are now required to keep child sitemaps on the same scheme, host and effective port, and malformed XML fails the warmup instead of reporting a successful partial run. Move a child sitemap outside that authority onto the index authority before upgrading.

Client-supplied forwarding aliases and rewrite hints such as Forwarded, CF-Connecting-IP, True-Client-IP, X-Original-URL and X-Rewrite-URL are now removed before the origin request. PRISM continues to set its own trusted X-Forwarded-For, X-Real-IP, X-Forwarded-Proto and X-Forwarded-Host.

Nothing needs reconfiguring. One thing changes what a cache in front of PRISM does, and it is the point of the release: a response PRISM declines to store — a 401, 403, 407, 429, or a temporary redirect the origin did not grant a lifetime — is now sent with no-store instead of public, max-age=3600. If something downstream was absorbing those, it will stop, and requests that were being answered at the edge will reach PRISM. That is the correct behaviour: a 403 pinned at a CDN is a page Googlebot cannot reach until the TTL expires.

Two further effects are improvements you may notice in metrics rather than behaviour. Cache hit rates on sites with tracking parameters will rise, because a URL whose query is entirely stripped now shares the bare URL's key instead of missing it. And the Age reported on a cache miss now includes the render, so it will read a few seconds higher than before — the same number a hit for the same document reports.

Nothing needs reconfiguring. Several things behave differently and all of them are visible from outside PRISM, so they are listed with what to expect rather than left to be discovered.

Redirects that call themselves temporary are no longer cached. A 302, 303 or 307 was kept for the route's TTL; now it is stored only if the origin states a lifetime for it. On a site that redirects a lot this means more renders and a lower hit rate — and correct behaviour, because a temporary redirect pinned for an hour outlives the thing it was temporary for. If yours are stable enough to cache, have the origin say so with Cache-Control: max-age=....

A write drops what was rendered for the same URL. A successful POST, PUT, PATCH or DELETE now invalidates the cached GET of the URL it addressed, on the replica that handled it. Expect a render on the next crawl of that page. This is local — a fleet still needs the durable purge journal, which is 1.8 work.

Cached entries start at their real age. The clock used to start when the entry was stored, which was after the render. It now accounts for the origin's Date, its Age, and the time the fetch and render took. For an origin that sends Date and no Age, this costs a few seconds off a lifetime measured in hours — not something you will see. For an origin already behind a CDN, the Age it reports is now respected rather than partly discarded, which is the point: PRISM stops handing a response a full lifetime it was never granted.

Errors and refusals are announced honestly. A cached 5xx is advertised for cache.error_ttl_secs rather than the route's TTL, and PRISM's own 429/5xx carry no-store. If a CDN in front was absorbing origin outages on the longer number, raise error_ttl_secs — it now governs both ends.

A new response header. Cache-Status (RFC 9211) is sent alongside X-Prism-Cache, which is unchanged. Nothing reads it unless you point something at it.

Two new settings, both defaulting to current behaviour: cache.honor_no_transform (false) and server.max_request_headers (100, the value already in force). Neither changes anything unless you set it.

Downstream caches will hold less than they did. Nothing needs changing to adopt this, but three things behave differently and all three are visible from outside PRISM:

  • A cached 5xx is announced with the lifetime it is actually stored under (error_ttl_secs, default 30s) rather than the route's. If you relied on a CDN in front to absorb origin outages for longer, raise cache.error_ttl_secs — that setting now governs both ends.
  • 401, 403, 407 and 429 from the origin are no longer stored unless the origin states a lifetime for them. If your origin returns one of these for pages you want cached, it must say so with Cache-Control.
  • PRISM's own 429/502/503/504 carry no-store. If a CDN was absorbing PRISM's rate limiting for you, that shielding now happens at PRISM's own limiter rather than downstream — check security.rate_limit_per_ip is where you want it.

No configuration keys were added, removed or changed, and no default moved.

One theme: a cache must not outlive the condition it cached. Three responses that describe a moment — an origin failing, an origin refusing, PRISM itself refusing — were each being kept, or advertised, for the lifetime of a page.

If you installed 1.5.1 from a package, the systemd unit it dropped caps memory below what the default pool needs. dpkg, rpm and apk treat /usr/lib/systemd/system/prism.service as a configuration-adjacent file, so check it after upgrading and make sure it reads MemoryMax=6G. Then systemctl daemon-reload && systemctl restart prism. Nothing else in this release needs action.

  • New: [bot.verify] — forward-confirmed reverse DNS for claimed crawlers. Off by default; every key defaults, so existing configs are unaffected. Turn enabled on to observe and enforce on to act.

  • One tightening can refuse startup. An admin.bearer_token that still reads as a placeholder (CHANGE-ME…) is rejected at load. Only deploy/k8s/config.toml ever shipped one; if you copied it, replace the value (openssl rand -hex 32).

  • The memory guard is stricter. Chrome is now costed at 525 MiB per tab rather than 375, calibrated to what the nightly scale matrix actually observes. A container limit that used to pass silently may now log a startup warning — that warning is the point, and the fix is more memory or fewer tabs, not a smaller estimate.

  • Cache keys escape | and % in their URL portion. Keys are in-memory only, so there is nothing to migrate; purge patterns are escaped the same way and keep matching. Patterns written against URLs containing a literal | need no change.

  • Kubernetes users: both shipped paths now require a licence Secret (prism-license) and the chart ships a NetworkPolicy that is on by default. Set networkPolicy.origin or your renders cannot reach the origin — the chart says so on install.

  • Every new config key defaults sanely: license.recheck_minutes = 60, license.grace_hours = 72, admin.metrics_bearer_token unset. No schema breaks.

  • One behaviour tightening can refuse startup: a [cache] vary listing a credential header (cookie, authorization, host, x-prism-*, x-forwarded-*) is now rejected at load. No correct deployment had one — such a value copied session tokens into cache keys and handed them to the renderer — but if yours did, remove the entry before upgrading.

  • GitHub Releases are gone as of this version; packages and Docker images are unchanged channels, and each release's SBOM + checksums now live at pkg.trident-cache.com/prism/provenance/<tag>/.

Nothing requires configuration changes. Three behaviours become visible:

  • Purges start working in viewport mode. Single-URL and pattern purges now remove the ::mobile/::desktop variants they always claimed to remove. If your workflow relied on purges being no-ops (it did not), purge volume is unchanged — but entries actually disappear now.
  • Warmup renders each URL once per device class when render.viewport is enabled — twice the warmup renders, so budget warmup windows accordingly. Admin /render reports the variants in a variants array; its top-level fields are unchanged for single-variant deployments.
  • Coalesced waiters wait out their leader instead of 503ing. Under load spikes, requests that previously failed at timeout+5s now wait as long as the leader may legitimately work. Tail latency replaces spurious 503s.

Six things change behaviour an operator has to plan for. Everything else in this release is a fix that needs nothing.

  1. Hit rates fall where the origin said not to cache. Pages carrying no-cache, max-age=0, s-maxage=0 or Vary: * are no longer stored at all, and a shorter s-maxage or max-age now caps PRISM's own TTL instead of being overruled by it. Magento sends no-cache on more routes than most operators expect.

    Check what your origin actually sends on the routes you care about before reading the new numbers as a regression. An origin that means a page to be cached should say so, and s-maxage is the directive aimed at PRISM specifically.

  2. Size security.rate_limit_per_ip before upgrading. Concurrent requests for an uncacheable URL each render now, where one render used to answer all of them. The default budget is three renders per IP per second, and behind a CDN every visitor resolves to the same edge address unless server.trusted_proxies lists your proxy — so one burst on a no-cache page can exhaust it for the whole site. Set trusted_proxies, raise the budget, or both.

  3. The admin API now authenticates. Anything calling /purge/*, /render or /warmup without a credential stops working, including scripts that relied on loopback being open. Set admin.bearer_token to a value you control; without one PRISM generates a token per run and logs it at startup, which is secure but changes on every restart.

    Probes are unaffected — /health, /ready and /startup never required a credential, so orchestrators and the Docker healthcheck need no change.

  4. admin.peers now requires a configured admin.bearer_token. Startup refuses the combination, because a generated token cannot be shared between replicas and every forwarded purge would be rejected. Set the same token on every replica.

  5. Subresources on hosts PRISM cannot resolve are now blocked. The guard that stops a rendered page reaching private addresses used to wave through any hostname it failed to resolve. A third-party asset whose DNS is slow or briefly down is now refused rather than left to fail on its own.

    The cost is bounded on purpose — an unresolved verdict is never cached, so a flaky lookup refuses the one subresource in front of it and nothing more. Only applies to cross-origin subresources, and only while security.block_private_cidrs is on; your own origin is exempt either way.

  6. X-Forwarded-Host is pinned when server.public_base_url is set. A client Host that does not match the configured one is replaced with it. If your origin legitimately serves several hostnames through one PRISM instance, leave public_base_url unset — behaviour is then unchanged.

Added​

  • GET /explain answers why PRISM did what it did with a URL. The question it exists for is the one the roadmap says has cost days of back-and-forth on live deployments: why will my page not cache? Until 1.6.0 the honest answer was a header value and a shrug, because the decision was an if-else chain producing string literals and nothing could report on it. The render decision and the cache decision now carry typed reasons, so this is mostly a matter of asking them.

    Three answers, asked separately. Routing: would this path be rendered, and if not, which condition stopped it — pass user_agent to ask as a specific crawler. Cache: what is stored right now, how old, how much longer it counts as fresh, whether grace may serve it, and every variant under the same canonical URL. Policy: a hypothesis — given a status and cache_control as your origin would send them, would the result be stored, for how long, what would be announced downstream, and if not stored, why not, in a sentence.

    It renders nothing, fetches nothing, and does not touch the cache it reports on. RenderCache::inspect reads without promoting the entry in the LRU or counting a hit — a diagnostic that moves the numbers it reports would raise the hit rate of anyone debugging a low one, and there is a test that fails if it ever starts to.

  • The release record is published where the documentation sends readers. versioning.md says the changelog is the release record and that each release's Upgrading section is the migration guide. The repository is internal, so that sent a reader to a file they could not open. The changelog is now a page on the documentation site, generated from CHANGELOG.md at build time rather than copied — a second hand-maintained changelog diverges within a release or two, and the divergence stays invisible until someone acts on the stale half.

  • cache.honor_no_transform, defaulting to false — and the default is the decision worth explaining. RFC 9111 §5.2.2.6 tells an intermediary not to transform a representation, which is exactly what replacing a shell with rendered HTML does. PRISM has always ignored it; that is now a stated choice with a switch and a counter rather than an omission.

    The directive exists to stop intermediaries the origin does not control — a carrier proxy recompressing images, a transcoding gateway. PRISM is not one: the origin's own operator installs it, points it at their own site, and pays for that transformation. Magento and other stacks also emit no-transform inside boilerplate Cache-Control with no intention about rendering.

    The failure modes settle it. Obeying by default means an upgrade silently stops rendering for someone who bought PRISM to render, surfacing weeks later as "our SEO stopped working" with no obvious cause. Not obeying breaks the letter of a directive in a way nobody observes. The second is the cheaper mistake — and prism_no_transform_ignored_total counts every response it applies to, whatever the setting, so an operator who disagrees can see it and flip one key.

    With it on, the response is served from the origin unchanged and marked x-prism-fallback: no-transform. no-transform says nothing about storing, and PRISM does not read it as if it did: an origin that sends it with a lifetime still gets its response cached.

  • server.max_request_headers, defaulting to 100 — which is what the HTTP layer would have used anyway. The value is not the point; depending on it without saying so was. PRISM's inbound limits were all inherited defaults, and an upgrade that changed one would have moved this deployment's posture with nothing in the diff to notice.

    The header bytes stay non-configurable and bounded at 8 KiB plus 4 KiB per header, answering 431. That same ceiling bounds the request target, which is why there is no URL-length setting: the URI is parsed out of the same buffer, so there does not need to be one.

    tests/inbound_limits.rs boots the real connection builder on a loopback port and checks what a socket refuses — 300 headers under a ceiling of 100, 60 under a configured 20, a two-megabyte request target — rather than asserting what a constant says. It needs no Chrome, no Docker and no origin, so it runs on every commit.

  • Cache-Status, the field RFC 9211 defines, sent alongside X-Prism-Cache rather than instead of it — that header predates the standard and deployments read it. The point is that a CDN, a browser devtools pane or an engineer reading someone else's stack can now see what PRISM did without knowing PRISM exists: prism; hit, prism; fwd=miss, prism; fwd=bypass, and a stale hit as prism; hit; detail=stale — a qualified hit rather than fwd=stale, because the reader was answered from the cache, which is exactly what serving stale is for.

    No ttl parameter: the RFC makes every parameter optional, and the remaining freshness is derivable from the Age and Cache-Control sent beside it — both of which had to be corrected in this same series before that was true. A third statement of the same number is one more thing that can disagree with the other two.

  • A write now drops what PRISM rendered for the same URL. RFC 9111 §4.4 requires a cache to invalidate the target URI when an unsafe method comes back with a non-error status. PRISM did not: a POST that changed a page passed through to the origin, and the rendered copy kept being served to crawlers until its TTL ran out, with nothing in the system aware the page had changed. Every variant is dropped — an exact-key delete would leave the mobile and Vary copies serving the old page while reporting success — and the URL is resolved by asking the render pipeline, exactly as the admin purge does, so the three ways an entry can be dropped cannot drift apart.

    Safety is read as RFC 9110 §9.2.1 defines it, a closed set of GET, HEAD, OPTIONS and TRACE, so an extension method counts as unsafe rather than slipping past an allowlist.

    Single-replica deployments are fully correct; a fleet is better than it was, not finished. A write landing on one replica leaves the others holding their own copies, and closing that needs the durable purge journal rather than another best-effort fan-out — see 1.8 in ROADMAP.md.

  • The nightly evidence suite keeps its evidence and says when it fails. It uploaded nothing and notified nobody: every number in the capacity model is sourced from a run's console output, which GitHub deletes on its own schedule — a citation with an expiry date — and a suite that runs at 02:47 and reports to no one is a suite nobody reads. Both failures so far were caught only because someone happened to be merging at the time. Each job now captures the k6 summary and PRISM's own logs as a 90-day artifact, and a reporting job opens a labelled issue on any scheduled failure and closes it when the suite is green again. A test pins the artifact retention, the issues: write permission, the failure condition and the close path — each of which is useless without the others.

  • Release verification now proves the shipped image renders. It checked that the image starts, proxies to the origin and answers /health — which is precisely what PRISM does when it has no licence, so every release so far was verified against behaviour a plain reverse proxy would also pass. The job now mounts a time-limited CI licence (PRISM_CI_LICENSE secret), requests the page as a crawler, and asserts both x-prism-rendered: true and the presence of content that exists only after JavaScript runs. The origin fixture had to grow up for this: at 92 bytes and titleless it could never satisfy content validation, which is on by default and wants a title, 100 characters of text and 1 KiB of HTML. Missing secret fails the job — an unverifiable release is not releasable.

  • A written versioning and deprecation policy (versioning in the docs): what each version-number component promises, which surfaces are versioned (config schema, admin API, metrics, packages — not internal Rust APIs), and the deprecation contract: nothing disappears within a major, deprecated keys warn at startup naming their replacement (the [detect] section already operates this way), removal only at the next major, and deny_unknown_fields stays strict throughout.

  • A shipped ops pack (deploy/ops/): a Grafana overview dashboard, 12 alerting rules in both plain Prometheus and PrometheusRule CRD form, a runbook per alert, and documented SLOs whose latency objectives anchor to what the nightly evidence suite already enforces. A structural test keeps it honest: every referenced metric must actually be exported, the two rule files must carry the same alert set, and every alert must have a runbook.

  • The licence grace window is now visible to metrics. prism_license_grace_active is 1 exactly while the window is open — the interval where "renew now" is still actionable. prism_license_valid stays 1 through grace (rendering continues), so without this gauge the first alertable signal was rendering having already stopped.

  • A published capacity model (deployment/capacity in the docs): the two-regime cost model (warm hits vs cold renders), the memory formula the binary itself enforces at startup, render-throughput and warmth-budget formulas with a worked example, and crawl-burst levers — every number cited to its source (nightly evidence runs, scale-matrix observations, or code constants) rather than asserted.

  • A Chromium update watch with a 7-day adoption SLA. The renderer's pin made builds reproducible and also guaranteed nothing updated it — the one component cargo-deny cannot see. The daily security-audit workflow now compares the pin against Debian's indexes (newest of bookworm-security and bookworm, compared with dpkg semantics, failing closed on unreadable indexes) and opens a labelled issue the day a newer chromium ships; SECURITY.md commits to adopting it within 7 days. A structural test keeps the watch's teeth in place.

  • Claimed crawlers can be verified with forward-confirmed reverse DNS. [bot.verify] checks a claimed Googlebot/bingbot/Applebot/Yandex/Baidu identity the way its operator documents (PTR into the operator's domain, forward-confirmed to the same IP), with per-(IP, identity) verdict caching. enabled observes — verdicts in logs and prism_bot_verification_total{outcome}; enforce treats a proven-spoofed crawler UA as a human — in bot-only mode that means proxied, never buying a render or a Chrome tab. Deliberately asymmetric: only a completed check that positively failed convicts — timeouts and resolver failures are unverifiable and never enforced against, because a degraded resolver must not demote real crawlers to the raw JavaScript shell. Off by default.

  • Credential headers are refused as [cache] vary dimensions. A value there becomes part of every cache key and is forwarded to Chrome, so vary = ["cookie"] meant session tokens copied into keys, handed to the renderer, and one cache entry per visitor — an eviction attack by configuration. There is no deployment where that was intended; startup now refuses it.

  • Data-handling and disaster-recovery documentation. Both answer procurement rows with facts already true and enforced by tests: the renderer is anonymous, credentials cannot reach it, the cache is memory-only and TTL-bounded (RPO zero — losing every replica loses warmth, not data), and the warmup is the recovery procedure.

  • The licence is re-evaluated at runtime, with a grace window and an in-place reload. It used to be evaluated exactly once, at load: expiry mid-run kept rendering until a restart, and the restart then silently stopped it — two cliffs, both surprises, the second one discovered from an SEO traffic drop. An hourly probe (config: license.recheck_minutes) now re-verifies the file without touching the live state; a renewal dropped in place is picked up within one interval with no restart; an expiry starts a loud grace window (license.grace_hours, default 72) that is screamed about on every tick before rendering actually stops. POST /license/reload applies a renewal immediately, and a reload that fails to verify keeps the previous state — a bad file must never un-license a running fleet.

  • Read-only metrics credential. admin.metrics_bearer_token is accepted only for GET /metrics, so Prometheus never holds a token that can purge the fleet; a test proves the scope opens exactly one door.

  • Audit events on every mutating admin call — endpoint, caller IP, parameters, entries removed, peer fan-out results — on the audit tracing target, so destructive actions stop being invisible to log review.

  • Licence visibility. /status reports version, license and licensed; /metrics gains prism_license_valid and prism_license_info{status=...} — the first alertable signal for the silent expiry-then-restart bypass.

  • Attributed rate-limit and purge metrics. prism_rate_limited_total{limiter="ip"|"domain"} answers which budget shed a request (the documented CDN 429 incident could not be told apart from domain limiting), and prism_peer_purge_failures_total is, until durable purge lands, the only sign a purge did not converge.

  • A Helm chart (deploy/helm/prism) mirroring the hardened Kubernetes manifest: same securityContext, probes and resource envelope, a ServiceMonitor wired to the scrape-only metrics token, and a chart that refuses to install without a real configuration rather than inventing one. appVersion is coupled to the crate version by a shipped-config test.

Changed​

  • The wire-level contract is asserted positively. The tests checked that a refused response did not carry max-age=3600, which a missing header or a different lifetime would also satisfy. They now require no-store exactly.

  • The test fixture builds on a supported runtime, reproducibly. It ran on Node 20 while react-router declares engines.node >= 22.22.0 — npm warned and carried on, so it worked by accident. It also ran npm install without copying the lockfile first, so the fixture the tests exercised was not the one the lockfile describes.

  • The wire-level test tier runs in CI. The hurl suite asserts what a socket receives and nothing executed it, which is how three of its assertions came to contradict a deliberate, unit-tested decision — that Vary names request headers, so X-Prism-Variant must not appear there — without anyone noticing. Switching it on caught two of the defects above the same day. A test guards that every compose stack under tests/e2e is referenced by CI, because a directory of tests nobody runs reads as coverage while being none.

  • scripts/check is the way to verify the tree. Asking cargo test for a count of failures answers 0 when the test binaries did not compile, because a build that produced no results produced no failures either — and piping discards the exit code that said so. Twice that became a confident "0 failures" while nothing had run, hiding a stale duplicate of a whole test module and then three files that no longer matched a struct they construct. The script asserts that tests ran and passed, refuses a run below a floor that would indicate a target silently not being built, and keeps cargo's exit status.

  • The documentation site's canonical URL points at the host that serves it. It named docs.trident-prism.com, which 301s to devdocs.trident-prism.com — so every page told search engines its canonical address was one redirect away from itself. Awkward anywhere, specific here: PRISM exists to control how crawlers see a page.

  • The render/bypass decision is a function, and its reason is a type. The choice was an inline if-else chain in the request handler with a string literal at each branch. The strings were already stable and already logged; what they could not do is be enumerated, so nothing could assert that the decision covers its own input space and nothing downstream could act on a reason without matching text.

    It is now decide_request(RequestFacts) -> (RequestDecision, DecisionReason) — no self, no request, nothing but the eight facts already computed by the time the choice is made. The logged strings are unchanged, so existing queries and dashboards keep working.

    What that buys is testing the whole space rather than the paths a test happens to walk: all 256 combinations of the eight facts across both modes, asserting properties instead of a table. Nothing renders without a licence whatever the other seven say. Bot-only never renders for a human — the property the privacy posture rests on. A render implies every condition was satisfied. Every reason is reachable, so no branch is dead code wearing a name. Making the licence check reachable-past fails the first of those by name.

  • Only the statuses a cache is allowed to store are stored. The rule that kept 401, 403, 407 and 429 out of the cache unless the origin said otherwise is now the rule RFC 9111 §3 actually states, asked as "is this on the permitted list" rather than "is this one of four".

    The case that showed it was the temporary redirect. A 302, 303 or 307 was kept for the configured hour — so a page moved aside for maintenance, an experiment or a geography stayed moved for crawlers long after it came back. A redirect that calls itself temporary is the one thing a cache should not pin. Their permanent counterparts, 301 and 308, are on the list and still store, as do 200, 203, 204, 206, 300, 404, 405, 410, 414 and 501.

    Server errors are exempt, and the exemption is where it belongs. A 5xx answers to cache.error_ttl_secs, which is the operator's explicit permission, so it does not pass through the §3 gate — applying it there would refuse every server error and switch off error caching entirely, the deliberate behaviour that shields a struggling origin. That regression was caught by an existing test rather than by review.

  • One cache decision, computed once. Storage lifetime was decided in the render pipeline while the client-facing Cache-Control was built separately in the proxy, and the storage rule itself was written twice — once for the foreground render, once for the background refresh. Four defects fixed earlier in this series were each an instance of the same shape: two places deciding, so they could differ. A 5xx kept for thirty seconds and announced for an hour; statuses stored that should not have been; an origin policy with no number left as a bare word; an entry whose clock started after the render that produced it.

    EffectiveCachePolicy now decides all of it at once — whether to store, why not when it will not, for how long, from what age, and whether the grace window may serve it. Both render paths call it; the duplicated rule is gone.

    The refusal reason is the new capability. A page that would not cache used to say so with a header value and nothing more, and answering why on a live deployment has taken days. There are four reasons, each a different conversation and each carrying both a stable token for metrics and a sentence for a person: the origin forbade shared caching; the origin granted a lifetime of zero, which is a permission without a duration rather than a prohibition; the status needs the origin's explicit permission; error caching is switched off.

    What guards it is an invariant test rather than a list of cases: across every status, origin directive and error-TTL combination, whatever is announced must equal whatever is kept. Reintroducing the original 5xx defect fails it with "kept for 30s and announced for 3600s".

  • A cached entry no longer starts life younger than it is — RFC 9111 §4.2.3. PRISM read the origin's Age header, ignored Date entirely, and started the entry's clock at the moment of storage. Both halves were wrong, and the second one especially for this product: the document arrives from the origin and then spends seconds inside Chrome before anything is stored. On the shipped 30-second error TTL a five-second render was a sixth of the lifetime, granted for free because nothing counted it.

    The corrected initial age is now the larger of two estimates, as the RFC requires: what Date implies, and the Age header plus the time the fetch and render actually took. An origin behind its own cache reports Age; an origin in front of nothing reports only Date; a response that passed through both is credited with whichever is more honest about its age. A Date ahead of our clock is treated as skew rather than as a response from the future, and an unreadable one as absent — the same reasoning Age already carried.

  • A cache policy without a number is completed rather than ignored. An origin sending Cache-Control: public has stated a policy and no lifetime. PRISM read the header's mere presence as a decision, suppressed its own synthesis, and stored the entry for its configured TTL anyway — an hour, by default. Downstream was left to derive freshness from ETag and Last-Modified, which rendered responses deliberately do not carry because the body was transformed, so it had nothing to work from at all. PRISM now keeps the origin's own words and appends the lifetime it is actually using: public becomes public, max-age=3600, s-maxage=3600. A header that already carries max-age or s-maxage is still left untouched, and one that forbids storing — no-store, private, no-cache — is never given a lifetime, because PRISM is not keeping it either.

  • A momentary refusal from the origin is no longer given the route's lifetime. 401, 403, 407 and 429 were stored under the configured TTL whenever the origin sent no directive of its own. None of them describes the page: they describe this request. A 429 raised for one second by an origin's rate limiter was kept for the configured hour and served to everyone after it, turning the origin's own protection into an outage that outlasted what it was protecting against. RFC 9111 §3 does not list these among the statuses a cache may store heuristically, and PRISM no longer does either — unless the origin states a lifetime for one, in which case it is honoured and capped as usual. skip_authenticated already keeps credentialed requests out of the cache; this covers the unauthenticated request that happens to receive one of these. Every other status is untouched, which has its own test.

  • PRISM's own error responses carry no-store. Its 429, 502, 503 and 504 are statements about the current instant — a limiter that is full now, an origin unreachable now. A CDN that stored one would go on refusing clients PRISM would already serve, and PRISM would never see those requests to know it was happening.

  • A cached error is no longer advertised for longer than it is kept. PRISM synthesises Cache-Control when the origin sent none, and that synthesis read the route's lifetime without looking at the response status. The render path meanwhile stores a server error under error_ttl_secs. With the shipped defaults the two disagreed by a factor of 120: a rendered 5xx was held for 30 seconds and announced as public, max-age=3600, s-maxage=3600, so a CDN or browser in front of PRISM kept an error page for an hour — long after the origin had recovered and long after PRISM itself had dropped it. The announced lifetime now matches the one actually used, and a 5xx that will not be stored at all (error_ttl_secs = 0) is announced no-store rather than offered downstream. Responses that are not server errors are unaffected.

    This changes what downstream caches do with your error pages. If you relied on the old behaviour to absorb origin outages at the CDN, raise cache.error_ttl_secs — that setting now controls both ends, which is what its name always implied.

    The decision moved out of the request handler into synthesized_cache_control, which had no test coverage of any kind before this change: the name appeared only in src/server.rs and src/config.rs. It now has seven, including the boundaries of the 5xx range and the case where the origin stated its own policy.

  • The memory guard is calibrated to the heavier of the two workloads we measure. CHROME_BYTES_PER_TAB came from the pool soak (375 MiB/tab on ordinary pages) while the nightly scale matrix settles the same 4-tab pool at 2.4–2.6 GiB on adversarial ones — so the guard stayed silent exactly when pages got heavy, and the shipped 8-tab default appeared to fit 4 GiB. Now 525 MiB/tab, derived from that observation, with a test that fails if the default pool ever silently fits a 4 GiB limit again.

  • The licence is commercial everywhere, and the crate can no longer be published. LICENSE was MIT, Cargo.toml said license = "MIT" and the README footer offered the project under it — while the product documentation and the binary itself say PRISM is commercial and refuses to render without a signed licence file. The repository was granting the rights we invoice for. LICENSE now states commercial terms conditioned on that licence file, Cargo.toml carries license-file plus publish = false (crates.io never deletes a published version, so one cargo publish would have been permanent and irreversible), and cargo-deny skips unpublishable crates while still checking every dependency. Package metadata declared MIT too — dpkg -s prism and rpm -qi prism printed it on every release so far — and pointed its homepage at the internal source repository, which answers 404 to every customer who follows it; both now say what is true. Since the software ships without source, the terms travel with it: packages and the container image carry LICENSE. A structural test keeps every surface in step.

  • GitHub Releases are retired; provenance moved where customers can reach it. The repository is internal, so release assets answered 404 to everyone a release is for — installation is the package repositories and the Docker registry, and always was. Each release now publishes its CycloneDX SBOM and SHA-256 checksum manifest at pkg.trident-cache.com/prism/provenance/<tag>/, staged, verified and promoted through the same fail-closed chain as the packages; structural tests pin the ordering and the production assertions. The publish phase drops its contents: write token (the Release was its only writer), the broken-by-design dist/install.sh is gone, and CHANGELOG.md is the release record.

  • The origin's cache directives now decide what PRISM keeps. Only no-store and private were read; everything else was passed to the client and ignored on the way in. A page marked no-cache was stored, reused without validation, and served alongside the very header saying that was not allowed — stale prices and stale availability, advertised as fresh.

    One limitation decides most of this: PRISM cannot revalidate. There is no conditional request back to the origin before a stored render is reused, so a directive asking for validation before reuse can only be honoured by not keeping the entry. That now covers no-cache, max-age=0, s-maxage=0, and Vary: *, which must never be reused for another request.

    Where the origin does grant freshness it caps what PRISM keeps rather than being overruled by it: s-maxage first, since it exists to address shared caches, then max-age. A page the origin called fresh for a minute is no longer served for the hour PRISM was configured for.

    Grace is part of that arithmetic rather than an exception to it. PRISM serves an entry for grace_period_secs past its TTL while re-rendering, which is exactly what must-revalidate forbids, so for those the TTL is pulled back until the whole window fits inside the granted freshness. If nothing is left, nothing is stored.

    If your hit rate drops, look at what your origin sends on those routes. An origin that means pages to be cached should say so; s-maxage is the directive aimed at PRISM specifically.

  • A shared cache may no longer stretch what the origin granted. s-maxage carries the semantics of proxy-revalidate (RFC 9111 §5.2.2.10), so grace — which serves an entry past its TTL while a re-render runs — must not apply to it, and did. must-revalidate with no stated lifetime was ignored on the same path. The entry now records the origin's refusal and grace skips it, which also fixed the first attempt at this: subtracting the grace window from the TTL refused anything whose freshness was shorter than the grace period, so an s-maxage=60 page went uncached under a five-minute grace.

  • Responses now carry Age, and inherit the one they arrived with. RFC 9111 §4 requires a cache to say how old a reused response is. PRISM said nothing, so a downstream CDN started its own freshness clock from zero and the same seconds were counted on every hop. Incoming Age is now subtracted from remaining freshness too: max-age=60 with Age: 50 has ten seconds left, not sixty, and Age: 90 arrived stale.

  • A render the origin refused to a shared cache is no longer handed to the requests waiting on it. Coalescing broadcast it to every concurrent waiter, reaching them as x-prism-cache: HIT — a hit on something never cached. Not storing it did not make it less shared.

    This costs renders. Four concurrent requests for a no-cache URL are now four renders rather than one, each charged to rate_limit_per_ip — three by default, and one address for all traffic behind a CDN unless trusted_proxies is set. Size that budget before upgrading if your origin sends no-cache under crawl.

  • An authenticated shopper is sent to the origin instead of rendered. In render-all, Chrome renders with no session — no cart, no account, no prices that depend on who is asking. skip_authenticated kept that result out of the cache, so one shopper's page never reached another, but the shopper who asked still received it: an empty cart on their own cart page.

    A request carrying an Authorization header or any cookie is now proxied to the origin. Anonymous visitors on the same route still render. Bot-only is unchanged, and setting skip_authenticated = false restores the old behaviour.

Fixed​

  • Sitemap warmup no longer refuses the origin it was pointed at. The private-address resolver added in 1.7.1 enforced regardless of block_private_cidrs = false and of the configured-origin exemption, so http://spa:3000/sitemap.xml failed wherever spa is a container name on a private range — most compose and Kubernetes deployments, and the shape the documentation uses. Literal-IP testing never showed it, because a literal is not a name. The URL check and the DNS check now ask the same question. A host that is not the origin still may not answer with a private address.

  • A coalescing waiter registers atomically. When a leader's client disconnected, every waiter woke at once, each found nothing in flight, and each registered — several rendered the same URL while the rest were told the tab pool was full, a 503 for a page that renders perfectly well. The recovery path uses the same atomic operation the normal path always did, and an in-flight entry is now removed only by the registration that created it: a disconnected leader's guard no longer deletes the waiter that replaced it.

  • A recycle drains the tabs that exist. It waited for the configured tab count from a semaphore that can hold fewer — a launch where one tab failed to open, or a forgotten permit, shrinks it. A 7-of-8 pool could never drain: it waited out the whole 35-second budget, logged "renders still in flight" when there were none, then restarted underneath the healthy renders it never waited for. Those died as crashes and fed the circuit breaker, so one tab lost at boot degraded every recycle afterwards.

  • A completed render no longer walks the process table. The browser-memory gauge was refreshed inside release(), which the request awaits before it can answer, so every render paid for a read_dir of /proc plus two file reads per process — the host's whole table under the shipped systemd unit, which sets no ProtectProc. The five-second sampler already keeps the gauge current.

  • Admin sitemap warmup resolves through the private-address guard. The URL check looked at the name, and the fetch that followed resolved that name again — nothing required the two answers to agree, so a host that answers public once and private on the next query reached an internal address from inside the render host. Checking twice does not close that; there is always another lookup after the last check. The client now resolves through the same classifier the render path uses, and refuses a name that fails, times out, returns nothing, or includes a private address.

  • A release proves the commit it publishes passed its tests. release.yml refused a tag that was not on main, which says the code was reviewed and merged — not that it worked. The automatic trigger fires only after the required jobs succeed, but a manual workflow_dispatch reached the publish phases without any of that, and every release so far has been a manual dispatch: green by habit rather than by construction. A published version number cannot be withdrawn here, so the check belongs before the publish.

  • A retry cannot double the wait the caller was promised. Each render attempt carries its own global budget and the retry started a second one from zero, so with the shipped ten-second timeout a request could wait two thirty-second budgets before falling back — in render-all, that is a person sitting through both. The retry now gets what is left of one total budget rather than a fresh one of its own.

  • Repeated origin headers no longer corrupt cache policy or panic a connection. CDP joins repeated fields with newlines. private followed by public, max-age=3600 was consequently stored as public; other combinations reached an infallible header conversion and closed the connection on every MISS and HIT. Values are split back into repeated fields, the effective policy combines every Cache-Control, and the final conversion fails safe to no-store.

  • An iframe can no longer become the page's status and headers. Document response events are restricted to the main frame, so an embedded 404 with no-store and noindex cannot turn a cacheable 200 page into that response.

  • Render-all keeps human traffic available when Chrome has no capacity. A full tab queue or open circuit now sends crawlers a cache-safe 503 with Retry-After, while JS-capable human clients receive the origin response as documented. Tab acquisition and pre-navigation CDP setup are also inside the global render lifecycle budget, so a wedged browser cannot bypass that cap.

  • Tracking parameters cannot poison their shared cache entry. Chrome now navigates the same normalised URL used by the cache key.

  • Turning synthesis off does not turn safety off. Cache-disabled and authenticated cache-bypass responses with no origin policy still emit no-store; a downstream cache is never invited to keep what PRISM kept nowhere.

  • Encoded URLs can be inspected, invalidated and purged. Administrative and unsafe-write paths now escape canonical URL prefixes exactly as cache-key construction does.

  • Server errors stop at cache.error_ttl_secs. The ordinary page grace window no longer extends a 5xx beyond its advertised error lifetime.

  • Page-supplied JavaScript statuses are range checked. Values outside 200–599 no longer become PRISM's final HTTP status.

  • Origin Date reaches corrected-age calculation in production. The tested Date branch is no longer cut off by the preserved-header filter.

  • /explain recognises variant-only entries, authenticated storage bypasses, effective preserved origin policy, and encoded cache keys.

  • Malformed sitemap XML is an error, not an empty or partial success, and a sitemap index cannot pivot recursive fetches outside its parent authority.

  • Turning the cache off now turns it off downstream too. With cache.enabled = false PRISM stored nothing and went on announcing public, max-age=3600, so an operator who disabled caching had a CDN in front keeping rendered pages for an hour — with nothing in PRISM's own configuration to explain where the stale copies came from. The documented meaning of the setting is that caching is off, and downstream is part of "off". An origin that stated its own policy still keeps it: disabling PRISM's cache is not an instruction to countermand the origin.

  • A credentialed request is no longer offered to a shared cache. cache.skip_authenticated sends a cookie-bearing request past PRISM's cache, because the body answers that requester. PRISM kept none of it and announced public, max-age=3600 anyway — the same failure skip_authenticated exists to prevent, moved one hop downstream. That path also reported the origin's raw Age rather than the corrected one, and now goes through the same decision as every other.

  • Sitemap numeric references and CDATA are read. &#38; and &#x26; are the decimal and hexadecimal spellings of &, and arrived as literal text — so warmup fetched ?x=1&#38;y=2, a URL that does not exist, and reported success for it. A <loc> wrapped in <![CDATA[...]]>, which is the other legal way to write an unescaped & and what several generators emit, produced no URL at all: in the warmup report a silently dropped page and a page nobody listed look identical.

  • A refusal explains itself in terms of the status it was asked about. /explain gave every status outside RFC 9111 §3's list the sentence "401, 403, 407 and 429 describe the request rather than the page" — including for a 302, sending the reader after an authentication problem that does not exist.

  • A sitemap URL was truncated at its first escaped character. XML requires & to be escaped, so ?a=1&amp;b=2 is what a correct sitemap generator writes — and the parser stopped at the ampersand, warming https://example.com/search?a=1 instead. Every URL with more than one query parameter was affected, and silently: warmup reported success for a page it had never fetched.

    The cause is that a <loc> does not arrive as one event. quick-xml splits the content at every entity reference — Text("?a=1"), GeneralRef("amp"), Text("b=2") — and the parser took the first fragment and stopped reading. It now accumulates until the closing tag, resolving the five predefined entities and passing anything else through as written, since a URL that 404s is visible in the warmup report while a dropped entry looks exactly like a page nobody listed.

    Present since the sitemap parser was written; found while migrating to quick-xml 0.42, by writing a test for behaviour the migration touched.

  • What PRISM refuses to store, it no longer invites a CDN to store. 1.6.0 stopped keeping a 302, 401, 403, 407 or 429 without the origin's permission — and went on announcing public, max-age=3600 on the way out. So PRISM declined to hold a 403 for a second while inviting Cloudflare, Fastly or Varnish to hold it for an hour. For a renderer whose job is how crawlers see a site, that is an outage with our name on it: a 403 pinned at the edge is a page Googlebot cannot reach until the TTL expires. It was also worse than before 1.6.0, which at least kept the two halves consistent.

    The announcement is now derived from the storage decision instead of computed beside it, and carried on RenderResult and the cache entry rather than recomputed by the proxy — so a hit and the miss that produced it cannot answer the same question differently. When storage is refused and the origin said nothing, the response carries no-store; when the origin said something, it is left exactly as sent.

  • A URL whose query parameters are all stripped now shares the bare URL's cache key. normalize_cache_key returned its input verbatim when there was no query to strip, skipping the normalisation every other path applied. With an origin configured as http://shop:80 — an explicit default port, which is legal — that produced two keys for one page, so a crawler arriving with ?utm_source=... never hit the entry warmed for the bare URL and paid for a full render every time. Silent: the cache appears to work, the hit rate is simply lower than it should be. Present since the function was written.

  • The age reported on a miss is the corrected one. A ten-second render was advertised downstream as zero seconds old, so a hit and the miss that produced it reported different ages for one document. Coalesced waiters had the same defect.

  • The shipped memory ceilings no longer make the binary warn about itself. 1.5.0 recalibrated Chrome's per-tab budget from 375 to 525 MiB against the nightly scale matrix, which moved the startup guard's estimate for the default eight-tab pool from 3400 to 4600 MiB. The ceilings shipped alongside it did not move. dist/prism.service — and therefore every deb, rpm and apk of 1.5.1 — kept MemoryMax=4G, which is 4096 MiB, so a fresh install warned at startup that Chrome would be OOM-killed under load and that a restart policy would turn that into a restart loop. The unit is now MemoryMax=6G; the Docker and Kubernetes ceilings were already 6, and the Kubernetes and Helm requests move from 4Gi to 5Gi so the scheduler reserves above the working set rather than below it. A test now parses MemoryMax out of the shipped unit, reads the pool size a default config actually produces, and runs the real guard against them — string-matching the shipped files against each other is what let this through, since they agreed with one another while both were wrong.

  • The leak gate asks whether Chrome's RSS plateaus, not whether it grew. repeated_large_renders_do_not_leak_resources compared a baseline against a final reading, and both attempts to make that comparison meaningful failed. Warming up harder made it worse, because the warm-up is itself enough to trigger a recycle. Skipping the comparison when the recycle counter moved during the run still allowed a recycle just before it: the nightly reported 1178 MiB growing to 2517 MiB while every other case in the same run sat between 2397 and 2546, which is a browser seconds old reaching the ordinary steady state. Two readings cannot separate warm-up from a leak, because both grow; what separates them is that warm-up flattens. RSS is now sampled after each of six rounds, restricted to the longest trailing run of rounds sharing a recycle counter, and the final round-to-round step must be under 256 MiB. The selection is a pure function with tests that run without Docker, including one built from the failing run's real series.

  • Nightly results are reported when the suite is re-run by hand. The reporter carried github.event_name == 'schedule'. The workflow has exactly two triggers, so that excluded the manual re-dispatch: a re-run that went green could not close the issue a failure had opened, and a re-run that failed could not update it. It looked like it worked — an issue opened on a scheduled failure and closed the next morning — but that close was done by hand while the job was skipped. Guarded by a test that checks only the jobs calling github.rest.issues, since selecting a soak tier from a dispatch input is a legitimate use of event_name in the same file.

  • The Chromium watch can close the urgent issue it opens. It opens two titles — "update available" and "no longer installable" — but both the retitle and the close lookups matched only the first, so the urgent one could be opened once and never found again. Both lookups now recognise both, with a test that counts the matchers on executable lines: the first version of that test counted the phrase, which also appears in the comments explaining it, and passed while both lookups were still wrong.

  • The chart pins its rollout strategy instead of taking Kubernetes' default. 25%/25% removes a quarter of the render capacity mid-rollout and brings the replacements back with a cold cache — a default nobody chose, while the manifest has pinned maxUnavailable: 0, maxSurge: 1 all along. values.yaml has claimed since it was written that its "defaults mirror deploy/k8s"; the only thing checking that was the version string. A coupling test now covers the policy's excluded ranges, the strategy on both sides, and the NOTES warning, and CI asserts them in the rendered output.

  • acquire now honours the one deadline it was given. It set a deadline, then looped — fail fast, wait out a recycle, try again — while each attempt started a fresh ACQUIRE_TIMEOUT of its own, so a call could take twice the bound its caller was promised. That matters beyond the wait: the render pipeline's coalescing deadline is computed from a model that says one timeout, so waiters could give up on a leader that was still inside its legitimate budget and turn a slow recycle into 503s for exactly the requests coalescing exists to absorb. Each attempt is now bounded by what remains, via a pure attempt_budget in the style of the pool's other decision functions, with a test that fails against the old behaviour.

  • The five-minute guide no longer hides the licence. It never mentioned licensing — zero occurrences — while instructing the reader to request a page as Googlebot and expect "fully rendered HTML with all JS-generated content". A fresh install has no licence, so PRISM proxies and renders nothing, and the very first thing every evaluator tried returned the raw shell from the page whose job is their first five minutes. The guide now lists the licence as a prerequisite, mounts it in the docker run, says plainly what an unlicensed install looks like (indistinguishable from a working one except for the rendering), and points at /status → licensed as the first thing to check. A test keeps it that way.

  • The Chromium watch could not open the issue it exists to open. The patch that taught it to distinguish "a newer version exists" from "the pin is no longer installable" was lost mid-edit: the issue body referenced an installable variable the compare step never defined, so the stale path would have thrown a ReferenceError and opened nothing. Both halves are now present and the file was read back to prove it.

  • Ten documentation claims that did not match the code. admin-api.md reported a warmup status (completed) the enum never emits (complete) and implied the mutating endpoints are open when no token is set — PRISM generates and logs one, and refuses a non-loopback bind without either a token or an explicit opt-out. README named a metric that does not exist (prism_render_duration_sum_ms; the export is prism_render_duration_ms_sum) and a test count from an earlier era. SECURITY.md recommended MemoryMax=2G against README's 4 GiB floor, and told Docker operators to set block_private_cidrs = false — the change a shipped-config test refuses, and unnecessary, since the configured origin is already exempt. versioning.md promised support for two minor series where SECURITY.md's table says the current one. capacity.md called the manifest's 4 Gi a limit (it is the request; the limit is 6 Gi) and described the shipped pool as the code default (4 tabs versus 8), and cited two measurements to a k6 run that reports neither — one is derived from PRISM's own request logs and now says so. metrics.md called itself a complete reference while omitting 18 of 42 families, and shipped competing alert definitions reusing the names in deploy/ops with different thresholds; the families are documented and the alerts now point at the tested set. Two tests keep the reference complete and single-sourced.

  • Documented configuration that refused to start. The CDN recipe in README and cdn.md recommended vary = ["x-prism-variant"] — rejected at startup by the credential-vary validation 1.4.2 itself added, and contradicting the warning two paragraphs above it, which correctly says X-Prism-Variant is a response header. The edge header is now x-device-class, a name the edge owns, with the collision explained. troubleshooting.md's two remedies for HTTP 429 put rate_limit_per_ip under [server] and rate_limit_per_domain under [render]; both live under [security], both are per second not per minute, and an operator copying either during an incident got a process that would not come up. The React and Angular presets dropped the required selector: prefix. A new test parses every TOML block in README, SECURITY.md and the documentation site through the real deserializer and validate() — 103 blocks, all starting — so the whole class stays closed rather than these four instances.

  • Both shipped Kubernetes paths now produce a deployment that renders. The raw manifest — which the chart's README calls the reference — carried no licence volume, no mount and no [license] section, so it could not render by construction; the Helm chart mounted the Secret as a directory at /etc/prism/license, leaving the file at /etc/prism/license/license.lic while every document says /etc/prism/license.lic. Both failures were silent: pods go Ready, metrics look healthy, and crawlers get the JavaScript shell. The manifest now mounts a required prism-license Secret and points its config at it, the chart uses subPath to land the file exactly where the docs promise, and an unlicensed install says so in NOTES rather than looking like a working one.

  • CI renders the deploy artefacts instead of only reading them. A new job runs helm lint, templates every mode the chart supports (licensed, unlicensed, autoscaling + PDB + ServiceMonitor), asserts the licence lands on the documented path, checks the chart still refuses a missing configuration, and parses the raw manifests as Kubernetes objects. The chart is the path we recommend and nothing had ever rendered it.

  • Chromium pin moved to 151.0.7922.173, which is a security update, and the watch now knows the difference between "behind" and "unbuildable". Debian's indexes carry only the current version of a package, so when 151.0.7922.173 superseded .137 the pinned version stopped existing and every image build began failing — CI, e2e and releases alike, within hours of the daily watch reporting the pin as current. The watch now reports two distinct conditions: a newer version exists (adopt within the 7-day SLA) and the pin is absent from every index (broken builds, fix now), with the issue title and body saying which.

  • Cache keys are injective again — a collision was an unauthenticated poisoning primitive. normalize_cache_key percent-decoded the query with query_pairs() and rebuilt it with format!("{k}={v}"), so the separators came back as structure: ?a=b%26c%3Dd (one parameter whose value contains & and =) and ?a=b&c=d (two parameters) produced the same key. Because the render fetches the raw URL and stores the result under that key, one GET for the encoded form parked the attacker's page on the victim's entry — no credentials, no headers, no timing. The query is now re-encoded when the key is rebuilt, and three tests pin it: the collision, the same trick aimed at the strip_query_params denylist, and idempotency (a key fed back through normalisation must not move, or purges stop matching what is stored). Found by the 1.4.2 audit, which reproduced it by execution rather than by reading.

  • A grace window now closes even over a deleted licence file. reload refuses to mutate state unless the file verifies — right for a bad renewal, wrong as the only tool for grace expiry: a licence file deleted or corrupted after expiry left the live state licensed forever, contradicting the documented "rendering stops when the window closes". The watcher now sheds the licence explicitly when reload cannot (trident-license gained mark_expired), with the behaviour pinned by tests on both sides.

  • Documentation claims brought in line with what the code enforces. An adversarial docs-vs-code audit found seven statements without backing: --version does not report licence state (the startup log does), packages ship license.path commented out rather than defaulted, the auto-generated admin token is deliberately printed once at startup (the docs claimed "never logged"), and disaster-recovery quoted timing numbers and an 8-tab reference pool that nothing measures or runs (the CI stacks use 4 tabs). Docs now state the enforced numbers and the real defaults.

  • The k6 scenarios' failure accounting matches their own semantics. The spoofed-bot scenario exists to provoke 429s, yet hung its threshold on k6's built-in http_req_failed, which counts every 4xx/5xx — it could never pass (first execution: 98.9% "failed" while 98.3% of checks succeeded). Three sibling scenarios carried the same latent mismatch and passed only because their rates happened to stay under threshold. Every scenario now excuses exactly the statuses its own checks accept (per request where bots and humans differ — a rate-limited human still counts as a failure), and the spoofed-bot scenario's own failed_requests metric finally has a threshold.

  • The k6 load stack can actually launch Chrome. The nightly evidence suite's first execution failed all five k6 jobs: the compose lacked the sandbox settings (SYS_ADMIN, apparmor=unconfined, 1g /dev/shm) every e2e compose already carried, so Chrome could not create user namespaces. Mirrored verbatim and proven locally end to end before shipping.

  • Single-URL and pattern purges removed nothing in viewport mode. Cache keys carry the device class as ::mobile/::desktop; both purge sweeps anchored on | alone, so a price-change purge matched zero entries cluster-wide, reported success, and stale prices served to Googlebot until TTL. The tests covering the sweeps fabricated a |mobile suffix that key construction never produces — they now build every key through CacheKey itself, and matching derives from the same DeviceClass::ALL the keys do.

  • Warmup filled a cache its intended reader never looked at. With viewports enabled, warmup and admin /render rendered only the desktop variant, so Googlebot Smartphone — the crawler mobile-first indexing sends — missed on 100% of warmed URLs. What "warmed" means now comes from the pipeline that owns the keying decision, and both callers iterate it.

  • Coalesced waiters gave up before their leader legitimately finished. The waiter deadline was timeout+5s; a leader may wait out a 60s tab acquire and spend the 3× global render budget, twice. The deadline is now derived from those same constants, so the two cannot drift. A leader that dies still wakes its waiters immediately.

  • Wedged-Chrome recovery waited behind its own readers. Restart takes the browser write lock while read guards held CDP calls bounded only by the transport timeout (sized above a whole render), so recovery took 1–2+ minutes of full-instance 503s after a ten-second detection. Every CDP call made under a read guard is now bounded at 5s — control-plane calls finish in milliseconds on a healthy browser. Replacement-tab creation retries with a fresh guard between attempts before a slot is ever given up: one timeout cannot tell a wedged Chrome from a healthy browser on a saturated host, and treating it as a permanent loss would have let load drain the pool.

  • HEAD to a large resource returned 502. The Content-Length early-reject fired for every method, and a HEAD's declared length describes a GET body the origin never sends. Error responses in the proxy path also no longer build bodies for HEAD; the method-blind error helper is deleted.

  • Malformed trusted_proxies entries were silently dropped, and a dropped entry is indistinguishable from an absent one — a whitespace typo reproduced the documented site-wide CDN 429. Broken entries are now named in a startup warning (a startup error from 1.5.0), the parser rejects impossible prefixes, and IPv4 entries now match the ::ffff: mapped peers a dual-stack listener reports — previously "trust nothing" on [::] binds.

  • Startup warnings finally reach a logger. validate() warned before the tracing subscriber existed, so the memory-budget OOM warning and the [detect] notice had never printed once. They are returned to main and logged after logging initializes.

  • The documented installer failed checksum verification on every release ever made — it saved the artifact under a renamed file and verified against the published name. Files keep their names end to end now.

  • deploy/k8s/deployment.yaml shipped v1.3.3 after 1.4.0 released. The manifest now carries the crate version and a test couples the two: bumping either alone fails the build.

  • The response direction strips the full RFC 9110 §7.6.1 hop-by-hop list, matching the request side.

  • A bypassed page now says what it varies on. In bot-only mode a URL has two representations: what a crawler is rendered and what a person is proxied. Only the rendered one declared the dimensions telling them apart, so a downstream cache could store the human copy against the bare URL and answer a crawler from it — a request PRISM never sees, and so cannot correct. Both representations now declare the same dimensions, which is what RFC 9111 §4.1 matches on. Routes that never render claim no variance, so edge caches are not fragmented for nothing.

  • Rendering no longer strips the origin's security headers. The preserved list carried three and dropped the rest, so a crawler fetching the rendered copy got no nosniff, no referrer policy and none of the cross-origin isolation headers — while a human on the bypass path, on the same URL, got all of them.

  • Every way out of the render path declares the same variance. Not being a bot was one of several: a response that was not HTML, and a render that failed and fell back, both went out declaring nothing.

  • Cache directives are read conservatively. The last duplicate value won, so a second max-age could widen what the first granted, and an unparseable value was dropped silently — which read as "the origin said nothing" and handed the response PRISM's full TTL. RFC 9111 §4.2.1 asks for the first occurrence, or for the response to be treated as stale.

  • A background refresh can revoke what it was refreshing. One that came back no-store only declined to write the new copy; the old one stayed inside its grace window and kept being served.

  • The 5xx error budget is bounded by the origin. error_ttl_secs says how long PRISM will hold an error, not how long it may outlast what the origin granted.

  • Headers named by Connection are removed before forwarding. RFC 9110 §7.6.1 requires an intermediary to drop them: they describe the hop ending here. PRISM removed a fixed list and never read what the sender actually named, in either direction.

  • TTL rules are matched against the path without its query. A request carrying a query got two TTLs — the render pipeline asked with the query attached, the Cache-Control header was built without it — so an entry could be stored under one and advertised with the other for the rest of its life.

  • purge_all no longer loses track of bytes. It cleared the entry map, released the lock, and only then zeroed the byte counter. A store landing in those two instructions added bytes to a total about to be overwritten, and its entry stayed in the map with nothing accounting for it. Eviction decides against that counter, so entries it cannot see never create eviction pressure.

Security​

  • Every Chrome option PRISM supplies now reaches Chrome. chromiumoxide adds the two leading dashes itself; PRISM supplied them too, producing four-dash options Chrome ignored. This included the internal-host resolver rules and the requested new headless mode. A Chrome-backed regression test now reads the real child process command line rather than the builder's debug output.

  • Percent-encoded excluded routes stay excluded. /%61dmin/secret can no longer evade the default /admin/** rule and reach the renderer. Invalid escapes, control characters and dot segments fail closed.

  • Forwarding identity and rewrite headers are allowlisted by construction. Vendor aliases supplied by the client can no longer override identity or routing in an origin that trusts them.

  • The deployment guide now states the actual browser boundary: CDP Fetch interception is defence in depth, while network/DNS egress controls are required to cover WebSocket, WebRTC, popup and out-of-process-frame traffic.

  • The sitemap size limit now prevents an allocation instead of reporting one. Warmup downloaded a sitemap in full and then compared its length against the 50 MB ceiling — the comment beside the check said as much, "post download size check". An origin that omits Content-Length and streams five gigabytes would have PRISM hold all five before the limit had anything to say. The body is now read as it arrives and refused at the first chunk that crosses the ceiling, so the buffer never exceeds it by more than the chunk in hand. The Content-Length check stays as a courtesy: it only helps against an origin that declares its size honestly, which is not the one this guards against.

  • The two open image-size advisories are assessed rather than left unexplained. GHSA-w3rx-r6r6-pgpr and GHSA-5p2g-fcmc-qvqq are infinite loops in its ICNS, JXL and HEIF parsers. It arrives transitively through @docusaurus/mdx-loader, runs only while building the documentation site, and appears in no binary, image or package PRISM ships; no fixed version exists upstream. Those parsers are entered only for those three formats and the documentation contains none, so a test now fails the build if such a file is ever added. Recorded in SECURITY.md, because an open high-severity alert with no written assessment reads the same as one nobody looked at.

  • The Helm chart ships the egress policy the raw manifest has had since 1.3.0. It shipped none, so the path the README recommends was the less defended one: no network backstop for the SSRF case PRISM's own filter cannot fully cover — a public hostname that resolves to a private address, checked for the top-level navigation but not for every subresource. On by default, with link-local (cloud metadata) and CGNAT excluded like the manifest. The origin rule has no default, because a guessed selector renders a policy that looks complete and blocks the one destination the product exists to reach; leaving it unset omits the rule and NOTES says so.

  • A URL can no longer forge the vary section of its own cache key. Vary names and values were escaped; the URL in front of them was not — so a request for /p|vary:accept-language=de and a request for /p under an accept-language vary dimension produced the same key, and whichever arrived second was served the other's HTML. Unauthenticated, and reachable on any deployment with [cache] vary configured. The URL portion now escapes | and % (and only those: ; and = cannot be misread, since vary begins only after a literal |vary:, and escaping them would rewrite every ordinary query string in every key). purge_pattern escapes the operator's pattern the same way, so patterns keep matching and globs are untouched.

  • The service-account token is no longer mounted into the pod that runs a browser. Neither Kubernetes path set automountServiceAccountToken: false, so a renderer compromise handed the attacker this pod's identity with the API server — for nothing, because PRISM never calls the Kubernetes API. Off in the manifest, off by default in the chart, guarded by a test.

  • A shipped placeholder credential is refused at startup. deploy/k8s/config.toml carries bearer_token = "CHANGE-ME-BEFORE-DEPLOYING" so the field is visible where an operator looks — but applying that tree unedited produced a running admin API whose token is printed in a public example, granting purge, warmup, render and licence reload. validate() now refuses any obvious placeholder and names the field. Omitting the key instead would be worse under Kubernetes, where each replica would generate a different token, so the file stays a template and the shipped-config test asserts it is refused for that reason.

  • The admin API is no longer unauthenticated on loopback. It purges the cache, forces renders and drives warmup, and it ran open whenever it was bound to loopback — a boundary that does not exist inside a container, where the address is shared with every sidecar and anything that gets a shell. PRISM now generates a bearer token for the run when none is configured and logs it.

    Set admin.bearer_token to keep it stable across restarts. Probes are unaffected: /health, /ready and /startup never required a credential, so orchestrators and the image's own healthcheck need nothing. insecure_no_auth is now the only way to run without authentication, and the startup refusal it used to override is gone — there is nothing left to refuse.

  • A redirect keeps the cache policy the origin gave it. The redirect path retained Location and dropped everything else, so a 302 marked no-store reached the storage decision with nothing to decide on — it was kept for PRISM's configured TTL and then advertised as public. Login, locale, canonical and retired-product redirects are exactly the ones an origin marks that way. Security headers on a redirect survive now too.

  • The origin's Age actually reaches the cache. The arithmetic for inheriting it shipped in this release; the header did not. Origin responses are filtered through a preserve list that did not mention age, so nothing ever reached the code that subtracts it, and every entry still started at zero. The age arithmetic is also saturating now — Age is a number an origin can get wrong, and the previous form panicked rather than clamping.

    Time spent rendering is still not counted: created_at starts when the entry is stored and the document arrived a render earlier. That is bounded by render.timeout_secs.

  • admin.peers without a configured token is refused at startup. A generated token cannot be shared, so each replica signed forwarded purges with its own and every peer rejected them — the fan-out silently did nothing and the other replicas served stale renders until TTL expiry, which is the failure peers exists to prevent. Set the same admin.bearer_token on every replica.

  • Liveness no longer mints a Chrome tab per call. /health is exempt from admin authentication so orchestrators need no credential, and the check it runs creates and closes a tab. A verdict is now reused for five seconds, which bounds a flood to one tab per window while staying well inside any probe interval. The cheap checks in front of it — crashed pool, no usable slots — are not cached and still answer immediately.

  • The DNS guard fails closed. It exists because PRISM cannot see what Chrome's own resolver will return, and it passed through every hostname it could not resolve — on the reasoning that the request would fail anyway. That holds when a name genuinely does not resolve, and not when it resolves for Chrome but not here, from a different cache or a different server, which is the case worth guarding against.

    An unresolved verdict is never cached, so the availability cost stays with the single request rather than taking the host out for the resolver TTL.

  • A forged Host no longer reaches the origin as X-Forwarded-Host. PRISM overwrites Host with the origin's own name, so the public one is passed along separately — and it was passed along exactly as the client sent it. A storefront deriving its base URL from that header would build canonical links, redirects and sitemap entries pointing at a host the visitor chose.

    Where public_base_url is set, the operator has said which host this instance serves, and anything else is replaced with it. Where it is not, the client's Host still travels: PRISM is then the edge as far as it knows, and refusing it would break every single-hop deployment.

Documentation​

  • The allowed_origins examples now match something. Every example on the security page was a URL, and the matcher compares bare hostnames — so each documented configuration matched nothing, and matching nothing blocks every subresource rather than failing loudly. An operator following the page got a rendered document with no CSS. A test now fails if an example regrows a scheme or an uppercase letter.

  • The published claims were checked against the code. "Zero unsafe code" was false — one unsafe block exists, a libc page-size lookup — "631 tests" was stale by four hundred, and "3 independent security audits" contradicted this repository in writing, which says internal review rounds do not count as independent. All three are corrected and a test now fails the build if any returns. The footer's Privacy Policy and Terms links pointed at #; they now point at the published documents.

  • rendering/status-codes.md claimed that not caching 5xx was the default when error_ttl_secs defaults to 30, so anyone who wanted errors uncached and trusted the page got the opposite. The settings table in configuration/cache.md is now parsed and compared against CacheConfig::default(), so the reference an operator reads cannot drift from the binary again.

  • ROADMAP.md was describing work that had already shipped in 1.5.3, and an external assessment scored cache correctness partly on that. It now records what is done, and sequences the route from the assessment's 7.5 to 9 as three releases — with the ceiling stated, since an independent penetration test, named customers and a second maintainer are not things code can deliver.

  • Five fixes that had been applied to one surface and not its twin: the per-tab figure in the capacity model, the #[ignore]d test count in README (173 → 223, measured), the "you may need to disable block_private_cidrs" advice that survived in troubleshooting.md after SECURITY.md dropped it, WarmupStatus::Failed which the admin API reference never listed, and the supported-versions table still naming 1.4.x as current.

  • The roadmap records Phase 1 as delivered, item by item with what makes each checkable, and names which of the six ranked deal-breakers it closes and which remain.

  • Three pages described behaviour this release changed. The README still said startup refuses an unauthenticated admin API on a non-loopback address; the admin page still said a null token leaves the API unauthenticated; and the cache page still said authenticated requests merely skip caching rather than bypassing rendering. All three now match the code, and the shared-token requirement for admin.peers is written down where an operator configuring replicas will meet it.

  • trusted_proxies is documented as what it does. Three places said an empty list trusts all proxies "for backward compatibility". The code has always meant the opposite, and its tests pin that. Following the documentation was expensive: behind a CDN every visitor resolves to the same edge address, and the per-IP render limit — three by default — then applies to that one address on behalf of the whole site.

CI​

  • Wire-level Hurl contracts and rendered deployment manifests are now release prerequisites for the tag-triggered release path.

Internal​

  • The 1.4.0 admin-auth default is pinned by a test that fails when it reverts. It lived only in untested main.rs wiring while the area's one test asserted a token-less server serves everything — reading as though open-by-default were the contract. The default path is now walked from PrismConfig::default() through the same resolution main uses; the old test remains as what it is, the explicit opt-out contract.

  • Production publication has teeth. An unscoped || true made the promotion ssh chain unfailable (every rsync could fail, step green) — in the very step hardened the day before. Production verification now runs under pipefail and asserts the released version in the apt and apk indexes before staging is deleted; reachability alone read a five-months-stale apk index as healthy. Docker :latest moves only after those assertions pass, not in the Package phase. A structural test pins all three against executable lines, not comments.

  • A provenance guard that could never have run. ci.yml refused to publish a tag not on main, and fetched with git fetch --no-tags --depth=0 — which git rejects outright as "depth 0 is not a positive number". The guard was added between releases, so its first execution was the 1.4.0 tag: it had never run once and failed where failing was most expensive. release.yml carries the same guard and fetched correctly, which is the only reason that tag was still publishable. A test now reads both copies.

  • The apk index is rebuilt when promoting to production. The production APKINDEX was last written in March, so every .apk published since — 1.1.0 through 1.4.0 — was uploaded and stayed invisible: apk add prism served 1.0.0 while 1.4.0 sat beside it on disk. The promotion rsync runs with --ignore-existing, which is right for package files and wrong for an index, and the metadata script is invoked with --skip-apk for production as well as staging, so nothing rebuilt it either.

    Staging was correct throughout, which is why release verification never caught it — that job installs from staging, the one repository whose index was current.

    The production index is still stale until it is rebuilt on the host once; this fixes the next release, not the current state.

  • Publishing fails closed. The staging upload, the metadata rebuild and the repository verification each carried continue-on-error: true, so a failure in any of them left the job green — and the GitHub Release ran on if: always() without depending on the repository verification at all. A release could therefore be announced against packages that were never delivered, which is what happened twice this week: the Docker image existed, the packages and the Release did not, and nothing failed.

    The release now waits for the repositories to be proven installable, and that proof asserts the installed version equals the tag rather than merely printing something. The package host's SSH key is pinned from a secret instead of being learned with ssh-keyscan on every run. Two tests read the workflow and fail if either gate is removed.

    This needs a new repository secret, PKG_DEPLOY_HOST_KEY, before the next release. Generate it once from a trusted network with ssh-keyscan -H <package-host>.

  • unknown-git is enforcing. The allowlist named both git dependencies and then only warned about anything else, which is a signal nobody had to act on.

  • CI runs the Rust gates when what compiles the Rust changes. The change filter matched **/*.rs, Cargo.toml and Cargo.lock, so a rust-toolchain.toml or deny.toml edit skipped every one of them — which is exactly backwards, since those decide which lints run and which dependencies are allowed.

  • read_body's error is boxed. Rust 1.98 widened clippy::result_large_err to catch it: the error variant is a whole Response, 144 bytes carried by every Result including the success path that never uses them. No behaviour change.

    Worth noting for whoever hits it next — CI pins its actions by SHA but resolves the Rust toolchain as stable, so a release of Rust can fail the build with no change to this repository. That is what happened here.

1.3.3 — 2026-08-19​

Fixed​

  • The internal origin survived in JSON-LD. PRISM replaced it in attributes — href, src, action, content, poster, data-src, srcset — and nowhere else. Structured data lives in the text of a script element, so none of that reached it, and a crawler was handed an entity identifying itself by a hostname that resolves nowhere:

    "url":"http://pwa-app:3000"
    "target":"http://pwa-app:3000/search?q={search_term}"

    It is the part of the page a search engine trusts most and the one part that was not rewritten. Blocks are now parsed, their URL-valued strings rewritten at any depth, and re-serialised.

    Two boundaries are deliberate and pinned down by tests: only strings that are the origin are touched, never one mentioned inside prose — editing inside a sentence would be rewriting the page's content — and only application/ld+json, because a string in ordinary application code may be a cache key or a comparison whose meaning would change.

    Parsed rather than string-replaced because a blind swap can leave a block unparseable, and a search engine discards a block it cannot parse: a corrupted rewrite is worse than the wrong host it was fixing. Anything that does not parse going in is returned untouched.

    Affects storefronts that derive JSON-LD URLs from window.location and set public_base_url. A Magento storefront building them from the store's configured base URL — most of them — was never affected.

1.3.2 — 2026-08-18​

The other half of the warmup bug: 1.3.1 stopped entries from being refused, this stops them from being unreachable.

Fixed​

  • Warmed cache entries could not be found by the traffic they were warmed for. Warmup rendered every URL and stored it, /status counted the entries, and every crawler request missed and rendered the page again — filing a second entry beside the first. Measured on the reporting host: warmup stored 5 of 5 URLs with no errors, and a Googlebot request for one of them was a MISS.

    Accept was folded into the cache key. It is appended to every request so Chrome negotiates with the origin the way the client would, which is a rendering concern; the same list also fed the key. Live traffic always sends Accept and looked under <url>|vary:accept=html, while warmup and the admin /render endpoint send no headers and filed their work under a bare <url>. Both sides worked and never met.

    The key now carries only headers listed in [cache] vary. Nothing is lost: a client that refuses HTML is turned away before the cache is consulted and a passthrough is never stored, so every entry holds the same rendered HTML whatever Accept asked for — the dimension split the cache three ways for identical content. An operator who lists accept still gets it, as a bounded class, and now gets it on requests that send no Accept too, which is what that class already meant.

    No configuration change is needed. If warmup appeared to do nothing on 1.3.1, this is why.

Internal​

  • Accept-Encoding behaviour is now covered by tests. It was already correct — one entry carries both the plain and gzip forms and the encoding is chosen per request, while the header is still forwarded to the origin on a passthrough — but nothing held it in place next to a change to how the cache key is built.

1.3.1 — 2026-08-18​

A cache that filled at random, and an HTTP/2 advisory that landed the day after 1.3.0 shipped.

Fixed​

  • A page's cache headers were taken from whatever it fetched, not from the page. Reported from the field as: warmup renders every URL, /status shows the entries, and live requests still miss. Measured against the affected origin, PRISM answered with no-store, no-cache, must-revalidate, max-age=0 for a document whose origin returns s-maxage=60 on forty consecutive requests — that header belonged to /graphql on the same origin, which is what PHP sends once session_start() runs.

    PRISM read the document's status and headers from wait_for_navigation_response(), which returns the main frame's most recently finished request. Any request the page makes overwrites that, with no filter for resource type, so a storefront calling its own API decided its own cacheability by coin toss. It now identifies the document by resource type from Network.responseReceived.

    Any headless storefront whose pages call a session-bearing API was affected — Magento with a JS frontend most of all. If you saw a cache that never filled despite successful renders, this was it; no configuration change is needed.

Security​

  • RUSTSEC-2026-0258 — h2 0.4.15 accepts unbounded empty DATA frames, so a peer can hold a stream open and force allocation without sending payload. Advisory published 2026-08-17, one day after 1.3.0. Updated to 0.4.16.

1.3.0 — 2026-08-17​

Security and correctness fixes from an independent review of 1.2.0, plus the proxy-body streaming change that makes this a minor rather than a patch release. Several items change behaviour; read Upgrading before rolling out.

Upgrading​

  • Vary no longer advertises X-Prism-Variant; it advertises User-Agent. Vary names request headers, and X-Prism-Variant is one PRISM writes onto the response — so caches looked for a header nothing sends, found it absent every time, and collapsed bot/human and desktop/mobile into one object. User-Agent is what actually varies the body. It also caches poorly, which is the honest cost of being correct: to get useful edge caching, normalise the User-Agent at your edge into one header and list it in [cache] vary. That path was documented but never worked, because nothing told PRISM to vary on the header the edge was injecting.

  • Origin redirects are now forwarded instead of followed. A 302 reached clients as 200 OK carrying the target's body under the original URL. Bots now receive the redirect. Expect one extra round trip per redirected URL, and expect previously-duplicated URLs to consolidate in search results.

  • Chrome no longer accepts invalid TLS certificates. If your origin serves HTTPS with a self-signed or expired certificate, renders will now fail. Trust the CA, or set render.accept_invalid_certs = true deliberately.

  • block_private_cidrs = true in the shipped Docker and Kubernetes configs. The configured origin was always exempt from this check, so it never needed to be off; with it off, a rendered page could reach any other private address, including PRISM's own unauthenticated admin API.

  • prism validate now rejects configurations it used to accept: a server.origin that is not an absolute http(s) URL, and warmup.concurrency = 0. Both started successfully before and failed later — the second by panicking the process on the first warmup.

Fixed​

  • The image's health check could wedge its host. It probed /health every ten seconds. /health does a real CDP round-trip, so it created and closed a browser page six times a minute in competition with actual renders — and once the browser was loaded enough that the probe outran its timeout, the daemon killed it. On current Docker and containerd a killed exec strands its state on the /run tmpfs: an <exec-id>.pid file plus -stdout/-stderr FIFOs, none of which are ever reaped. That made the probe self-reinforcing, and one host accumulated 198k pid files — a 4 KB tmpfs page each, so the entire 776 MB of /run — with 396k FIFOs beside them holding 41% of the inode table. containerd could then not write init state and no container on the machine could start, including ones unrelated to PRISM. It presents as no space left on device on a host whose disk has plenty of room, because the full filesystem is a tmpfs. Now probes /ready, which reads atomics and cannot hang, every 30 seconds with a 60-second start period. A probe that completes normally leaks nothing, so this closes the leak rather than slowing it; the deployment docs carry the host-side cleanup for hosts that already have a backlog.

  • Accept-Encoding: gzip;q=0 was answered with a gzip body. The header was matched by substring, so an explicit refusal read as acceptance. notgzip matched, GZIP did not, and * did not.

  • Cache byte accounting was updated outside the lock that decides eviction. Two writers racing on one key could subtract an entry's bytes before they were added, wrapping the unsigned total to ~usize::MAX — after which eviction emptied the cache on every store, silently in release builds.

  • Purges reached one pod behind a headless Kubernetes Service instead of all of them. Each peer is now resolved and purged at every address, so one Service name covers an autoscaled fleet.

  • The internal origin leaked into rendered output. Anything a page derived from window.location — a JS-generated canonical, an og:url — named the internal address and was served to crawlers as the site's identity.

  • The Kubernetes manifest referenced a registry this project does not publish to with a digest of all zeroes, so applying it produced ImagePullBackOff.

  • The origin-timeout log reported "30s" regardless of proxy_timeout_secs.

Changed​

  • Proxy bodies stream instead of buffering. The pass-through path held each request and response body in memory so headers could be rewritten — headers that live in parts and never needed the body at all. At the shipped ceilings that was up to 10 MiB of request and 50 MiB of response per connection against a 10,000-connection limit, with nothing bounding the aggregate.

    Every limit that mattered is still enforced. Content-Length is refused up front in both directions, and Limited caps a body that under-declares its length or omits it entirely.

    Behaviour change: server.body_read_timeout_secs becomes an idle timeout — the longest a body may stall between frames — rather than a total deadline on reading it. That is the bound streaming wants: a total deadline cannot tell a slow attacker from a large legitimate upload, so raising it restores slowloris and lowering it severs a real upload over a poor link. A size violation now also surfaces mid-stream rather than before any bytes are sent; the statuses are unchanged (413 too large, 408 stalled).

    This is why the release is 1.3.0 and not a patch: it rewrites the request hot path and changes documented behaviour, and neither belongs in a patch version that operators would reasonably take without reading.

  • render.auto_scroll (default true) turns off the lazy-load scroll pass. On a tall page it spends its full five-second budget on every render, which sites that hydrate up front were paying for nothing.

  • render.accept_invalid_certs (default false) is the TLS escape hatch.

  • Test builds verify licences against a separate throwaway key (test-license-key feature), so no production-valid credential exists in the repository.

  • The release is now gated on the corpus, scale-smoke and coverage jobs, which describe themselves as required but were not in needs. cross is pinned by commit rather than tracking its default branch. CI cancels superseded pull-request runs.

1.2.0 — 2026-08-10​

Upgrading​

Four changes alter behaviour on upgrade even though nothing breaks at the API level. Read these before rolling out.

  • Memory: 4 GiB minimum for the default eight-tab pool, up from a documented 2 GiB. Four warm tabs measure ~1.5 GiB before any large render, so the old figure was below the working set of half the default pool. Anyone running to the previous recommendation is under-provisioned and will discover it under load. Lower render.pool.tabs before lowering the limit.

  • Vary: Accept is now emitted. PRISM varies its own cache on the representation class, so a CDN must be told. This changes cache fragmentation for every deployment behind one — expect a lower hit rate until the edge re-warms.

  • Non-HTML routes proxy the origin instead of returning 503. Correct — there was nothing to render — but an uncached feed now reaches the origin about three times per bot request where it previously produced a cheap error. Watch prism_passthrough_not_html_total after upgrading.

  • render.max_dom_nodes defaults to 400,000. A new rejection path that did not exist before. Documents above it are refused before serialisation.

Known limitations​

  • A non-HTML passthrough costs roughly three origin requests per client request.
  • render.timeout_secs above 30 requires the patched chromiumoxide pinned in Cargo.toml. Building against the published crate silently reimposes a 30-second ceiling on every CDP command; tests/cdp_timeout_e2e.rs fails if that happens.
  • Chrome intermittently stalls a single navigation and is retried, not fixed. On roughly one CI run in four, one route's Page.navigate receives no answer — no response, no events, no CDP traffic at all — while the browser stays healthy and the next request renders the same fixture in about two seconds. It affects a different local fixture each time, with no crash, recycle or circuit event in the run. The cause is not known. PRISM now discards the tab and retries once, which recovers the request; watch prism_render_retries_total, because a climbing counter means the stall is spreading and the mitigation is hiding it.
  • Not yet evidenced: commerce-scale validation, a Rendertron comparison, and a sustained soak. Suitable for a controlled rollout on that basis.

Added​

  • cache.honor_request_no_cache (default false). When enabled, a request carrying Cache-Control: no-cache, Cache-Control: max-age=0 or Pragma: no-cache is re-rendered instead of being served a stored response, per RFC 9111 §5.2.1.4, and the fresh render is stored so it refreshes the entry every later request sees.

    It is off by default on purpose: PRISM is a shared cache usually reachable from the public internet, and a revalidation costs a full Chrome render, so honouring the header unconditionally is a cache-busting denial of service that needs no credentials. Enable it only where callers are trusted, or behind a CDN that strips the header from client requests.

Added​

  • prism_render_retries_total — renders retried after a stalled navigation. One request costing two tabs, counted separately from prism_render_total so the mitigation above cannot hide its own cost.

Fixed​

  • A stalled navigation is retried once on a fresh tab. Chrome occasionally stops answering on one tab while remaining healthy; the tab was already being discarded, so a second attempt gets a clean one and generally succeeds. Only a timeout is retried — a systemic fault would get twice the load, a capacity refusal would steal a queue slot, and a page-local verdict would re-derive the same answer. See Known limitations: this recovers from the stall, it does not explain it.

  • A client that refused HTML was served a warmed cache entry. The Accept parser ranked quality weight above specificity, so text/html;q=0, */*;q=1 — a client asking for anything but HTML — was read as accepting it at q=1. RFC 9110 §12.5.1 is the other way round: the most specific matching range decides, and the broader ones speak only for the types the client did not name.

    The cache is what made it serious. Such a request was classed html, the same class an ordinary browser warms, so on a warm entry it was handed a stored render with no origin contact and no render at all — nothing on the render path was in a position to notice, and a cold-cache test passes against it.

    Consulting application/xhtml+xml was the other half: PRISM answers Content-Type: text/html and nothing else, so asking whether a client accepts a type PRISM never sends is how */* came to override a refusal that named text/html outright. Both the acceptability check and the cache class now ask about the one type PRISM emits, so they cannot disagree.

    Behaviour change: an Accept naming only application/xhtml+xml is now passed through rather than rendered. Browsers list text/html alongside it, so no real client is affected.

  • render.timeout_secs above 30 did nothing. chromiumoxide 0.9.1 takes both the per-command future's delay and the frame manager's navigation deadline from a crate constant rather than the configured request_timeout, so any page slower than 30 seconds failed as an opaque Navigation("Request timed out") regardless of configuration — and the bundled test stacks asking for 60 and 180 seconds were describing a bound that could not exist.

    PRISM now builds against a fork of v0.9.1 carrying one commit that threads the configured value through to both sites, pinned by revision in Cargo.toml. The fork point's tree was verified byte-identical to the published crate.

    This also explains a run of intermittent E2E failures previously written off as CI flakiness: always one route, always just past 30 seconds, always green on rerun, because the page usually finished inside the ceiling.

  • The runtime image installed a floating Chromium. Both base images are digest-pinned, but apt-get install chromium resolved against a live index, so rebuilding the same commit could install a different browser — the one component whose behaviour PRISM's serialisation, timing and size preflight depend on most directly. Now pinned by version via a build argument.

  • Accept was matched by substring, not parsed. text/html;q=0, application/json;q=1 contains the text "text/html" while explicitly refusing it, so PRISM read that as acceptance, rewrote the header it forwarded to text/html, and returned the rendered page the client had just rejected. Accept: text/plain was likewise upgraded to */*.

    There is now a real parse: media ranges with quality weights, q=0 treated as the refusal RFC 9110 defines rather than a low preference, the most specific matching range deciding the weight, and media-type parameters not mistaken for weights. Two decisions that were one are now separate — the client's own header is forwarded to the origin unchanged so negotiation stays theirs, while the cache is keyed on a bounded three-value class so the key cannot be inflated. A request that will not accept HTML is refused before a tab is taken.

  • A background refresh of non-HTML content counted two passthroughs. Both do_render and the background arm recorded it, for a refresh that served no client at all. Only the foreground path counts now.

  • Successful passthroughs were filed as size rejections. They incremented prism_oversize_rejections_total{reason="not_html"}, reporting a ceiling problem where none existed. That label is gone; the passthrough counter records the outcome once, where a client is actually served.

  • Forwarding Accept to Chrome made it a cache-busting dimension. The previous release forwarded the client's raw header so Chrome and the origin fallback would negotiate the same way. It was also keyed on verbatim, and the header is client-controlled and effectively unlimited — text/html;q=0.9,*/*; q=0.8,x=<nonce> is a valid HTML-compatible request — so anyone could force a fresh Chrome render per request and evict the cache on the way.

    The cache key is now a bounded three-value representation class rather than the raw header, keeping cardinality fixed while still distinguishing negotiations that produce genuinely different documents. The header sent to the origin is the client's own, unchanged — see the parsing entry above, which supersedes an earlier version of this fix that forwarded the class and so asked origins for representations clients had not requested.

    Accept is also declared in the response Vary. PRISM varies its cache on it, so a CDN not told about it would collapse those variants and serve HTML to a JSON client — the same defect already fixed once for configured dimensions.

  • Non-HTML passthroughs were still counted as errors in two places. The authenticated-skip branch recorded requests_error, and background refresh recorded a page-local render failure. Both now go through one is_client_success predicate rather than an inline check per call site, which is what let the miscount reappear twice after being fixed once.

  • A non-HTML passthrough was counted as an error and could serve the wrong representation. Refusing to render a non-HTML document is correct, but the request that follows is a success: the client receives the origin's own response with its own status. It was being recorded as requests_error and as a page-local render failure, so a site with one busy feed read as broken on every dashboard. Passthroughs now have their own neutral counter, prism_passthrough_not_html_total.

    The two origin requests a passthrough makes — Chrome's navigation, then the fallback — were also not equivalent. Chrome received only the configured cache dimensions, so a content-negotiated endpoint could hand it JSON while the fallback, which forwards the client's real headers, was given HTML and served it unrendered. Accept is now always forwarded to Chrome, whether or not it is a configured dimension, so both requests see the same representation.

    Known cost, measured rather than assumed: one passthrough currently makes three origin requests. prism_passthrough_not_html_total exists so that volume is visible, and the test asserts the count stays small and bounded.

  • Documents are now bounded before they are fully materialised, and the bound cannot be walked around. render.max_html_bytes was enforced only after Chrome had built the whole document, serialised it, pushed it across CDP and Rust had allocated its own copy — the peak the limit exists to prevent, paid three times before anything checked it.

    A preflight now walks the DOM inside the page and adds up what serialising each node would cost, stopping the moment the budget is exceeded and never concatenating: a string longer than the remaining budget is over by inspection and is never read, so scanning stays bounded even on a hostile page. Alongside it, render.max_dom_nodes (default 400,000) caps element count.

    Getting that estimate right took several passes, each closing a way past it:

    • attributes, comments, tag names and escaping expansion are all priced, not just visible text — a page whose weight is one 40 MiB attribute scored as almost nothing under the first version;
    • UTF-8 is counted exactly, including surrogate pairs, and unpaired surrogates as U+FFFD, so an emoji is four bytes and a lone surrogate three;
    • element names are counted in UTF-8 too, since Chrome accepts non-ASCII names and ten million two-byte characters is 20 MB in each of the opening and closing tags;
    • <template> contents are walked as additional roots — they live in a separate DocumentFragment that no walker rooted at documentElement enters, while outerHTML emits every byte of them;
    • processing instructions are priced, and CDATA sections as the escaped text the HTML serialiser actually emits — verified against a live render, which produced &amp; per character and no <![CDATA[ wrapper;
    • a node type with no rule of its own refuses the page rather than being guessed at, because the traversal's whole value is knowing what the serialiser will write, and for an unanticipated type it does not;
    • documents that are not text/html are not rendered at all. Every rule here describes the HTML serialiser; XML and XHTML serialise differently, and the same traversal would mis-price them in the under-counting direction. Such a response is proxied from the origin unchanged rather than refused — there was nothing to render, and the origin already produced the right answer.
    • comments and raw-text elements (<script>, <style>) are charged verbatim rather than with escaping expansion, which had inflated them enough to refuse documents comfortably inside the budget;
    • a preflight that cannot complete — an evaluation error, an unreadable result, a timeout — refuses the page instead of falling through to serialising it, since the traversal runs inside the page and a hostile document can reach that path deliberately.

    The exact byte check after extraction remains authoritative; the preflight exists to avoid reaching it. prism_oversize_rejections_total{reason} records which check refused each document, so it is visible if the cheap ones stop earning their place.

  • Three nanoid advisories were fixable and had not been fixed. 3.3.16 in tests/e2e/spa-fixture, about/website and demo/magento-pwa/pwa-app, all upgraded to 3.3.18. The accepted-advisories note previously claimed every open alert resolved to image-size; that was written after auditing only two of the projects and is corrected.

  • Shipped memory guidance disagreed with itself and with the default pool. The README recommended 2 GiB in three places while the packaged systemd unit set 4G and Kubernetes requested 3Gi — for a default of eight tabs, where four warm tabs alone measure about 1.5 GiB before any large render arrives. A limit below the working set is an OOM kill during the first spike, and the figure an operator copies is whichever file they opened. Everything now states 4 GiB, with the measurement basis and how to budget from real traffic, and a test fails if the surfaces drift apart again.

  • The Chrome memory gauge could report a process that no longer existed. prism_chrome_rss_bytes was written once per completed render, just before a possible recycle, and never afterwards — so following a restart it kept the outgoing browser's peak while the new one was a fraction of the size, and a shut-down pool kept reporting its last live reading. A gauge is read as "now"; this one could be minutes old with nothing to say so.

    It is now sampled every five seconds by a supervised task, refreshed immediately after a restart, and zeroed on shutdown. The new prism_chrome_rss_age_seconds makes staleness visible instead of implicit, and the HELP text states that the value covers the browser process and all its descendants.

  • A page that could not render opened the circuit breaker for every other page. Each render failure was fed to the shared Chrome circuit regardless of what it meant, so a document rejected on purpose by render.max_html_bytes — requested repeatedly, which is what a crawler does — tripped the breaker and stopped rendering site-wide while Chrome was in perfect health.

    Failures are now classified: systemic (the browser or its CDP connection is broken), capacity (no slot, or maintenance in progress), page_local (this URL cannot render) and policy (refused by configuration). Only systemic feeds the circuit. The rest are counted in the new prism_render_failures_total{class} so being circuit-neutral does not make them invisible, and a page-local failure returns the half-open probe instead of stranding it, which would stall recovery.

    Verified against a live stack: twelve consecutive 33 MiB rejections against a failure threshold of five leave the circuit closed, all twelve counted as page_local, and a healthy URL still renders. Killing Chrome outright is still detected and still recovers.

  • Requests waiting out a browser recycle were not counted against queue_max. Waiting rather than refusing is correct — a recycle is planned maintenance and the replacement is seconds away — but a waiter is still a queued request: it holds a client connection and will take a tab as soon as one exists. Queue depth was accounted per attempt, so a parked waiter left the count and the queue reported itself empty while an unbounded number of them accumulated. queue_max, the only bound between a slow recycle and running out of memory, applied to none of them. Queue admission now spans the whole acquire, waits included.

  • A planned browser recycle was counted and recovered from as a crash, and returned 503 while it ran. Every exit reached the monitor through one untyped channel, so closing Chrome deliberately — which is what a recycle is — was indistinguishable from it dying. The recovery restart then closed a browser again, emitting another exit, which was counted again: five recycles produced five crashes and five restarts, with 503s throughout, and the crash counter said the browser was unstable when nothing had gone wrong.

    Exits now carry the generation they belong to and why they happened. An exit is a crash only when it is the current generation, unexpected, and the pool is still running — a planned close is marked before it is made, and an exit from a generation that is already gone cannot restart the browser that replaced it.

    Requests arriving mid-recycle now wait for the new browser instead of being refused. A recycle is planned maintenance and the replacement is seconds away; refusing during that window turned routine maintenance into an outage on every eighth render at the shipped max_renders_per_browser. The wait is bounded by the deadline a queued request already has, so a recycle that never finishes still fails rather than hanging.

    Measured before: 5 recycles, 5 crashes, 5 restarts, 503s. After: 20 renders across 2 recycles, 2 restarts, zero crashes, zero 503s, circuit closed.

  • A request body that arrived too slowly was answered 413 Payload Too Large. Slow and oversized are different client faults; the read timeout now answers 408 Request Timeout. Telling a slow client to send less does not help it, and conflating the two hid slow-client attacks among genuine size rejections in the logs.

  • The admin listener accepted connections without bound. It serves liveness, readiness, metrics and purge, so a caller holding connections open could starve the health signal an orchestrator depends on. Capped at 64 concurrent connections; past that, connections are refused rather than queued, because an admin request that cannot be served promptly is worse than one that fails fast — the caller is usually a probe with its own deadline.

  • POST /purge/pattern left device and vary variants behind. The glob was matched against the whole cache key, which ends in |mobile or |vary:..., so a pattern naming a page — the shape an operator writes to purge one URL — matched only the desktop, no-vary entry. The response reported success while every other representation kept serving the content that had just been invalidated. /purge/url already removed all variants; the pattern endpoint now does the same, and patterns written against a suffix still work.

  • Cache dimensions PRISM keys on were missing from the response Vary. With [cache] vary configured, PRISM fragmented its own cache correctly but advertised only X-Prism-Variant and Accept-Encoding. A CDN in front of PRISM was therefore free to collapse those variants and serve one language's render to another language's request. The configured dimensions are now merged into Vary alongside the origin's and PRISM's own.

Testing​

  • The scale suite asserted the wrong recycle metric and is now required in CI. prism_tab_recycles_total counts per-render page and context replacement, which happens on every release regardless of browser lifecycle, so the recycle test passed on ordinary traffic. It now asserts on prism_chrome_recycles_total, requires restarts to equal recycles (more means a planned close was recovered from as a crash), requires the crash counter not to move, requires the circuit to stay closed, and requires the memory sample to be fresh after the restart. settle() fails instead of continuing when renders never drain. The smoke tier runs as a required job; the full matrix stays on its own schedule.

  • The shared route corpus is now executed, and checked against the fixture. tests/corpus/routes.toml declared in its own header that four suites consumed it while nothing in the repository referenced it. It had drifted accordingly: /products/localized and /consent-variant were not fixture routes at all (/localized and /consent are), the soft-404 entry pointed at a path that only reached the catch-all, seven pages were marked "not yet implemented" long after they were built, and two of the product prices it asserted did not match what the fixture renders — never noticed, because nothing ever compared them.

    tests/corpus_manifest.rs now checks the manifest against the fixture's router and data on every cargo test, in both directions, with no server required. tests/corpus_e2e.rs drives every declared route through a real PRISM and checks the declared expectations. Both run in a required CI job.

  • Device-aware rendering is exercised by the e2e stack. [render.viewport] was absent from the e2e config, so every request rendered as bot-desktop and the fixture's mobile page served its desktop branch to a phone. The old assertion accepted the substring View, which both branches satisfy, so the suite passed with the feature switched off.

1.1.5​

Supersedes 1.1.4, which was tagged but never published: the release pipeline checks out the tag it is releasing, so a tag cut before a packaging fix cannot carry it. This is the first tag containing both the rpm signing fix and the CI change that stops Chrome's container limitations failing a package check.

Packaging and documentation fixes. The Rust crate is unchanged apart from the version, but two of the three ways of installing PRISM did not work.

Fixed​

  • dnf install prism failed on every RHEL, Rocky, Fedora and CentOS host. The repository publisher signed repomd.xml but never the packages. repo_gpgcheck validates the index; gpgcheck — which the shipped repository config sets — validates each package, and an unsigned rpm fails it with "Signature verification failed" however well signed the metadata is. Packages are now signed with rpm --addsign.

  • apk add prism failed on Alpine. The repository line was written as <url>/apk/<arch>, and apk appends the architecture itself, so it fetched .../apk/x86_64/x86_64/APKINDEX.tar.gz and reported "no such package" even though the index was published and reachable.

  • The documented nginx configuration did not strip the bypass header it said it stripped. proxy_set_header directives are not merged across levels: a location defining even one of its own discards every one from the enclosing server block. The example placed the strip at server level and proxy_set_header Connection "" inside location /, so a client could forge X-Prism-Bypass, and Host, X-Real-IP, X-Forwarded-For and X-Forwarded-Proto were all dropped as well.

Added​

  • dist/reverse-proxy/ ships complete nginx and Apache configurations, so they can be copied rather than transcribed out of a documentation page — which is how the proxy_set_header placement bug survived review.

  • dist/integration-testing/run-distro-matrix.sh runs the release's package verification locally against real packages. The matrix could previously only be exercised by cutting a release.

CI​

  • Package verification covers CentOS Stream 9, openSUSE Tumbleweed and Ubuntu 20.04, installs the arm64 packages under emulation, and runs the static tarball on alpine, debian-slim and busybox.
  • Those tests now start PRISM — proxy, admin endpoint, a rendered bot request and a clean shutdown — rather than checking that files were laid down.
  • The repository verification jobs install with gpgcheck=1 instead of 0, so they exercise what a user's machine does. With it disabled they passed throughout the outage above.
  • PRISM builds and passes its browser-free tests on FreeBSD 14.

1.1.3​

A review pass over the request path, plus the security work that came out of enabling code scanning. Most fixes here are paths that worked normally and failed on an abnormal exit — a panic, a cancelled future, a shutdown — or a check that recognised one spelling of a thing rather than the thing. No behaviour, configuration or interface changes; a 1.0.0 config still loads.

Security​

  • Alternative encodings of internal addresses were not blocked. The private-address check understood ::ffff:127.0.0.1 but not ::127.0.0.1. Both name the loopback interface; the check used to_ipv4_mapped, which by definition ignores the IPv4-compatible form. The same gap applied to 6to4 (2002:7f00:1::) and the well-known NAT64 prefix (64:ff9b::7f00:1), which wrap an IPv4 address inside an IPv6 one — including 169.254.169.254.

    This runs on DNS results as well as URL literals, so a hostname resolving to any of these spellings previously passed both the rebinding defence and subresource filtering. The embedded address is now decoded and checked, so 64:ff9b::808:808 still reaches 8.8.8.8 rather than the prefix being refused outright. Missing on the IPv4 side as well: broadcast, multicast, the 240/4 reserved block, and 0.0.0.0/8 — only the single address 0.0.0.0 was caught, not the range.

  • A self-referential origin was only caught in four spellings. The check that stops PRISM rendering itself compared the origin host against localhost, 127.0.0.1, ::1 and 0.0.0.0 as text. Binding to 0.0.0.0 makes the whole 127/8 range reach the process, so http://127.0.0.2:4000 against a 0.0.0.0:4000 listener was an infinite render loop that validation accepted. The host is now parsed as an address.

  • The benchmark harness returned exception text — including stack and local paths — to clients on a failed render.

Fixed​

  • A panic in a background re-render took that URL out of service. The foreground path has used an RAII guard since the coalescing rewrite; the background path cleaned up by calling remove_inflight on each of its six exit branches, and a panic skips all six. The leak does not heal: the URL then looks permanently in flight, so it is never re-rendered, and once past grace every miss waits out timeout_secs + 5 and returns 503 — until the process restarts.

  • Two clicks on /warmup started two warmups. The running check and the state reset were separate critical sections, so concurrent callers all observed a non-running state and all proceeded. Eight concurrent callers all won, crawling the same sitemap and interleaving their counters through one shared state. A panic in the sitemap fetch or parse additionally stranded the status on Running for the life of the process, refusing every later warmup while reporting one in progress.

  • A client hanging up mid-cleanup leaked a tab and its context. release moved the page and context out of the TabGuard before its first await, making that guard's Drop a no-op for the rest of the function. Chrome kept both until the next browser recycle.

  • Shutdown stopped accepting connections but not requests. The drain broke the accept loop and waited on the in-flight counter. An established keep-alive connection is just an open socket, and HTTP/1.1 lets the client send another request — which the server answered, for the whole drain window. Connections now shut down gracefully. Readiness already failed before the drain, so the hole was requests arriving on connections already handed over.

Performance​

  • Post-processing parsed the document once per transform. Each enabled transform ran as its own rewrite_str, so a page with everything enabled was parsed six times. Measured on a 1.9 MB document: 32.4 ms before, 16.7 ms after, on every cache miss. Not the 5.6x the parse count suggests — the handlers cost real time themselves.

  • Content validation copied the document to look at it. extract_title allocated a full lowercased duplicate to find one tag; strip_html_tags allocated a String for every <; the visible-text check built a Vec of every word and joined it into a new String to use only its length. Measured on a 1.46 MB page: 3.33 ms before, 1.80 ms after, per render.

  • A rewrite failure in post-processing was silent: the input was returned unchanged, so a page whose transforms did not run was served to bots with the scripts this exists to strip. It now warns.

Testing​

  • IPv6 is covered end to end rather than only as a predicate over IpAddr: the address cases run through URL parsing, which rewrites these hosts ([::127.0.0.1] normalises to [::7f00:1]), and a new test drives the real binary bound to [::1]. Nothing previously bound an IPv6 socket anywhere.

  • The keep-alive shutdown test was rewritten after it raced on CI. Two further attempts passed with the fix reverted — once because the process exited before the assertion, once because the request under test hit a stalled origin and "no reply" reads like "refused". It now runs an origin that stalls one path and answers another.

Security tooling​

  • CodeQL analyses Rust, the JavaScript projects and the Actions workflows. Secret scanning runs with push protection and generic-pattern detection. main requires a pull request with passing checks; v* tags cannot be moved or deleted. Every workflow declares permissions: explicitly and the release phases run with a read-only token.

  • Eight CodeQL findings fixed in the demo and e2e fixtures: incomplete single-pass HTML stripping replaced with a scan, and ten committed build artifacts removed.

1.1.2​

Fixed​

  • CDP requests now time out on PRISM's schedule, not the library's. chromiumoxide defaults every CDP request to a hard-coded 30s that owes nothing to render.timeout_secs. Any configuration above ~10s therefore had a ceiling it did not ask for: a slow navigation failed as an opaque "Request timed out" at 30s, before PRISM's own global budget applied, and short-circuiting the handling that closes and discards a possibly-wedged tab. It surfaced as intermittent render failures under load.

    The CDP request timeout is now derived from the global render budget, so PRISM's own bounds always decide first. It is never sized below chromiumoxide's default, so short configurations are unaffected.

    Operators who raised render.timeout_secs above 30 were silently not getting it. They are now.

Changed​

  • The demo app moved off react-router-dom to react-router v8. The former has no v8 and pins react-router@7, which made the patched version unreachable. Not shipped in PRISM; listed for completeness.

1.1.1​

Dependency and supply-chain maintenance. No changes to PRISM's behaviour, configuration, or public interfaces — the Rust crate is untouched apart from the version bump.

Security​

Dependabot alerts and automated security updates were enabled on the repository, which surfaced 77 advisories (2 critical, 32 high). None were in Rust — cargo audit reports zero vulnerabilities across the dependency tree, and the shipped binary and container image were never affected. All were in four JavaScript sub-projects that ship with the repo but not with PRISM: the docs site, the marketing site, the e2e render fixture, and the demo app.

Open advisories reduced from 77 to a small remainder, both criticals cleared (websocket-driver, shell-quote). Fixes came via Dependabot PRs, npm overrides for transitive packages it could not reach, and two migrations: the demo to react-router-dom 7, and the e2e fixture to react-router v8.

CI​

  • Build jobs for the four JavaScript projects. None had any CI, so every dependency bump against them was unverifiable — a broken build would only have surfaced on deploy. Two of these sites are publicly deployed.
  • E2E Docker jobs now run when their own inputs change. They declared needs: build, and build only runs for Rust changes, so a skipped dependency silently skipped all three jobs — meaning changes under tests/e2e/** or to a Dockerfile never ran the tests that exist to cover them. They build their own image and never used that artifact.
  • Dependabot extended to the four npm projects; it previously watched only the repository root for cargo, actions and docker.
  • tests/shipped_configs.rs loads every config file the repo ships, so a validation change that invalidates one fails in seconds locally instead of after a full multi-arch release build.

Fixed​

Several upgrades are only valid as a pair, and Dependabot proposes each half separately. Merging them individually broke all three JavaScript apps:

  • about/website: react 19 landed without next 16 (next 14 peers on react ^18). Completing the pair then required moving the Google Fonts @import above the @tailwind directives — Next 16's Turbopack enforces the CSS spec where webpack did not — and pinning TypeScript to 6, since 7 is the native rewrite and no longer exposes the compiler API Next type-checks with.
  • demo: react-dom 19 without react 19, and @vitejs/plugin-react 6 (peers on vite ^8) without vite 8. Apollo Client 4 additionally moved its React bindings out of the root entry, so seven files needed repointing at @apollo/client/react.
  • devdocs: the Docusaurus 3.10 bump requires @docusaurus/faster as an explicit dependency when future.v4 is enabled; 3.9 bundled it.
  • tests/e2e/spa-fixture: react 19 without react-dom 19.

1.1.0​

Config schema is backward compatible — a 1.0.0 config loads unchanged, since no keys were removed and every new key has a default. Several defaults changed behaviour, though, so read "Changed defaults" before upgrading.

Changed defaults​

Setting1.0.01.1.0Effect
security.rate_limit_per_ip0 (off)3Renders above 3/s from one IP get 429. Charged on cache misses only, so cached traffic — including everything behind a CDN sharing one source IP — is unaffected. Set to 0 to restore the old behaviour.
render.pool.max_browser_age_secs— (no recycling)3600Chrome is recycled hourly. Each recycle refuses renders for ~370 ms.
render.pool.max_renders_per_browser— (no recycling)5000As above, whichever comes first.
render.circuit_breakerconsecutive failures only+ failure-rate windowThe breaker can now open on sustained partial failure. Previously a renderer failing 75% of requests never tripped it.
Origin allowancehost + portscheme + host + portorigin = "http://x:8080" no longer also permits https://x:8080.
systemd MemoryMax2G4GSized for the shipped render.pool.tabs = 8. PRISM now warns at startup if the pool cannot fit the detected memory limit.

Fixed​

  • Chrome crash recovery never fired. The only crash signal was the CDP handler stream ending, which never happens while the pool holds the browser handle — so a dead Chrome was never detected and the pool never recovered. Replaced with an active liveness watchdog; detection now ~13 s. It also catches a browser that is running but no longer servicing CDP.
  • Cache-key collision across variants. Vary values were joined without escaping, so a client could craft a header value that collided with a different variant's key and be served that variant's HTML. Components are now escaped.
  • Transient accept() errors killed the process. EMFILE under load, or a client aborting during the TCP handshake, terminated the accept loop and exited. Now classified: retry, back off, or stop only for an unusable listener.
  • Panic on an origin-controlled header. A Vary response header containing a control byte panicked the connection task.
  • utm_* did not strip every matching parameter. The matcher used shell glob semantics where * stops at /, so utm_a/b survived into the cache key.
  • Circuit breaker: a single success no longer force-closes an open circuit.
  • HEAD error responses no longer carry a body (RFC 9110).
  • Page-supplied status codes (status_from_meta) are clamped to 200–599.
  • purge_pattern no longer holds the cache lock across its whole scan.

Added​

  • Cross-replica purge. admin.peers fans POST /purge/* out to sibling replicas, so one purge invalidates the fleet. Forwarded purges are never re-forwarded, so peers may list each other.
  • Readiness and startup probes. /ready and /startup answer from atomics (safe to poll often); /health keeps doing a real CDP round-trip and is the liveness signal. Readiness fails at the start of the drain so a load balancer can route away before requests are cut off.
  • Correlation IDs. x-request-id is accepted from upstream (validated) or generated, recorded on every tracing span, and returned on every response.
  • Subresource DNS-rebinding defence. Resolution is now validated for every subresource a page requests, not just the navigation URL.
  • Metrics: queue-wait histogram, rejections by reason, render phase timings, Chrome process-tree RSS, background-task failures, proactive recycles.
  • Background tasks are supervised — one dying is now logged and counted rather than silently removing the capability it provided.
  • Chrome restart backoff is jittered, so replicas failing together do not retry in lockstep.
  • Kubernetes manifests (deploy/k8s/) with probes, PDB, HPA on queue depth, and an egress NetworkPolicy.
  • Benchmark suite (bench/): self-contained comparison against Rendertron and a Puppeteer baseline.

Performance​

  • Cache hits no longer re-compress. Compression happens once at store time; a hit is a refcount bump. Measured 3.07 ms → effectively zero for a 1.1 MB page.
  • Route matching 1896 ns → 48 ns per request by classifying patterns once instead of running every glob.
  • URL security checks parse the URL once instead of twice (335 → 182 ns).
  • Tab release no longer serializes on an exclusive browser lock.

Dependencies​

  • chromiumoxide 0.7 → 0.9, which also removed a duplicated reqwest stack
  • lol_html 2.7 → 3.0, rand 0.8 → 0.10, lru 0.16 → 0.18, toml 0.8 → 1.1
  • Total dependency count 317 → 298
  • Base images pinned by digest; builds use --locked

1.0.0​

Initial release.