Skip to main content

Prometheus Metrics

PRISM exports metrics in Prometheus text exposition format on the admin API at GET /metrics.

Request Metrics​

MetricTypeDescription
prism_requests_total{status="hit|miss|bypass|stale|error"}counterTotal requests by cache status
prism_inflight_requestsgaugeCurrently processing requests
prism_rate_limit_rejections_totalcounterRequests rejected by rate limiter (429)

Render Metrics​

MetricTypeDescription
prism_render_activegaugeCurrently active renders
prism_render_queue_depthgaugeRequests waiting for a Chrome tab
prism_render_totalcounterTotal renders completed
prism_render_duration_mshistogramRender duration in milliseconds
prism_shadow_renders_totalcounterShadow renders completed (shadow mode)
prism_content_validation_failures_totalcounterRenders that failed content validation

prism_render_duration_ms histogram buckets: 100, 250, 500, 1000, 2500, 5000, 10000, 30000

Cache Metrics​

MetricTypeDescription
prism_cache_entriesgaugeNumber of entries in the render cache
prism_cache_memory_bytesgaugeMemory used by cached entries
prism_cache_evictions_totalcounterCache entries evicted (LRU or TTL)

Chrome Tab Pool Metrics​

MetricTypeDescription
prism_tab_pool_sizegaugeTotal tabs in the pool
prism_tab_pool_availablegaugeIdle tabs available for rendering
prism_tab_crashes_totalcounterTab crashes detected
prism_tab_creates_totalcounterNew tabs created
prism_tab_recycles_totalcounterTabs recycled after max renders
prism_chrome_restarts_totalcounterFull Chrome process restarts

Origin Metrics​

MetricTypeDescription
prism_origin_requests_totalcounterRequests forwarded to the origin server
prism_origin_duration_sum_mscounterCumulative origin request duration (ms)

Warmup Metrics​

MetricTypeDescription
prism_warmup_urls_processedcounterURLs successfully warmed up
prism_warmup_urls_failedcounterURLs that failed during warmup

Circuit Breaker​

MetricTypeDescription
prism_circuit_breaker_stategauge0 = closed (healthy), 1 = open (tripped), 2 = half-open (probing)

System Metrics​

MetricTypeDescription
prism_uptime_secondsgaugeSeconds since PRISM started
process_resident_memory_bytesgaugeResident memory (RSS) of the PRISM process
process_virtual_memory_bytesgaugeVirtual memory of the PRISM process
process_cpu_seconds_totalcounterTotal CPU time consumed
process_open_fdsgaugeOpen file descriptors
process_max_fdsgaugeMaximum file descriptors (ulimit)
process_threadsgaugeNumber of OS threads

Renderer Internals​

MetricTypeDescription
prism_render_queue_wait_mshistogramTime a render waited for a free tab. Distinct from render duration: this is the queue, that is the work.
prism_render_retries_totalcounterRenders retried after a stalled navigation. One request costing two tabs, so a climb here means the stall is getting more frequent.
prism_render_failures_total{class}counterRender failures by what they say about the shared browser. Only class="systemic" feeds the circuit breaker.
prism_navigation_duration_sum_mscounterSum of Chrome navigation time — the origin plus its JavaScript.
prism_html_extract_duration_sum_mscounterSum of DOM serialisation time. Rising against flat navigation means pages got larger, not slower.
prism_oversize_rejections_total{check}counterDocuments refused for size, labelled by which check caught them.
prism_passthrough_not_html_totalcounterRequests proxied unchanged because the origin served non-HTML. Not a failure, but each one cost a Chrome navigation first.
prism_requests_rejected_total{reason}counterRequests shed, by reason: queue_full, circuit_open, rate_limited, render_failed.
prism_rate_limited_total{limiter}counterRate-limit rejections attributed to the ip or domain budget. Which one is shedding is usually the whole question.

Chrome Process​

MetricTypeDescription
prism_chrome_rss_bytesgaugeResident memory of the browser process and all its descendants. process_resident_memory_bytes covers PRISM only, and Chrome is where growth happens.
prism_chrome_rss_age_secondsgaugeAge of the RSS sample above. Exported so a stale gauge reads as stale rather than as a current measurement.
prism_chrome_recycles_totalcounterProactive recycles on age or render budget — routine hygiene, unlike prism_chrome_restarts_total, which counts failures.

Licence​

MetricTypeDescription
prism_license_validgauge1 when the licence allows rendering. Stays 1 throughout the grace window.
prism_license_grace_activegauge1 exactly while the grace window is open — the interval where renewing still helps. The only metric that shows it.
prism_license_info{status}gaugeLicence state as a label: valid, trial, expired, no_license, invalid.

Bots and Fleet​

MetricTypeDescription
prism_bot_verification_total{outcome}counterFCrDNS checks of claimed crawlers: verified, spoofed, unverifiable. Rising spoofed is an attack becoming visible; rising unverifiable is your resolver degrading.
prism_peer_purge_failures_totalcounterPurge fan-outs that failed to reach a replica. Until durable purge lands, the only sign a purge did not converge.
prism_no_transform_ignored_totalcounterResponses rendered although the origin sent Cache-Control: no-transform. Counted whether or not cache.honor_no_transform is set, so ignoring the directive is a visible decision rather than an omission. Non-zero on a deployment that cares about the letter of RFC 9111 §5.2.2.6 is the signal to turn that setting on.
prism_background_task_failures_totalcounterSupervisor tasks that exited unexpectedly. Any non-zero value means a capability (crash recovery, orphan cleanup, limiter sweeps) is silently gone.

Scrape Configuration​

Add PRISM to your Prometheus scrape_configs:

scrape_configs:
- job_name: prism
scrape_interval: 15s
static_configs:
- targets: ['127.0.0.1:4001']
# If bearer_token is configured:
authorization:
type: Bearer
credentials: your-secret-token

Grafana Dashboard​

Import this JSON snippet as a Grafana dashboard to get started with key panels:

{
"title": "Trident PRISM",
"panels": [
{
"title": "Cache Hit Rate",
"type": "stat",
"targets": [
{
"expr": "rate(prism_requests_total{status=\"hit\"}[5m]) / (rate(prism_requests_total{status=\"hit\"}[5m]) + rate(prism_requests_total{status=\"miss\"}[5m]))",
"legendFormat": "Hit Rate"
}
],
"fieldConfig": {
"defaults": { "unit": "percentunit", "thresholds": { "steps": [
{ "value": 0, "color": "red" },
{ "value": 0.7, "color": "yellow" },
{ "value": 0.9, "color": "green" }
]}}
}
},
{
"title": "Render Duration (p50 / p95 / p99)",
"type": "timeseries",
"targets": [
{
"expr": "histogram_quantile(0.50, rate(prism_render_duration_ms_bucket[5m]))",
"legendFormat": "p50"
},
{
"expr": "histogram_quantile(0.95, rate(prism_render_duration_ms_bucket[5m]))",
"legendFormat": "p95"
},
{
"expr": "histogram_quantile(0.99, rate(prism_render_duration_ms_bucket[5m]))",
"legendFormat": "p99"
}
],
"fieldConfig": { "defaults": { "unit": "ms" } }
},
{
"title": "Active Renders & Queue Depth",
"type": "timeseries",
"targets": [
{ "expr": "prism_render_active", "legendFormat": "Active" },
{ "expr": "prism_render_queue_depth", "legendFormat": "Queued" }
]
},
{
"title": "Tab Pool",
"type": "timeseries",
"targets": [
{ "expr": "prism_tab_pool_size", "legendFormat": "Total" },
{ "expr": "prism_tab_pool_available", "legendFormat": "Available" }
]
},
{
"title": "Cache Entries & Memory",
"type": "timeseries",
"targets": [
{ "expr": "prism_cache_entries", "legendFormat": "Entries" },
{ "expr": "prism_cache_memory_bytes", "legendFormat": "Memory (bytes)" }
]
},
{
"title": "Circuit Breaker State",
"type": "stat",
"targets": [
{ "expr": "prism_circuit_breaker_state", "legendFormat": "State" }
],
"fieldConfig": {
"defaults": {
"mappings": [
{ "type": "value", "options": { "0": { "text": "CLOSED", "color": "green" } } },
{ "type": "value", "options": { "1": { "text": "OPEN", "color": "red" } } },
{ "type": "value", "options": { "2": { "text": "HALF-OPEN", "color": "yellow" } } }
]
}
}
},
{
"title": "Origin Latency (avg)",
"type": "timeseries",
"targets": [
{
"expr": "rate(prism_origin_duration_sum_ms[5m]) / rate(prism_origin_requests_total[5m])",
"legendFormat": "Avg Origin Latency (ms)"
}
],
"fieldConfig": { "defaults": { "unit": "ms" } }
},
{
"title": "Process Memory",
"type": "timeseries",
"targets": [
{ "expr": "process_resident_memory_bytes", "legendFormat": "RSS" },
{ "expr": "process_virtual_memory_bytes", "legendFormat": "Virtual" }
],
"fieldConfig": { "defaults": { "unit": "bytes" } }
}
]
}

Alerts​

The alerting rules ship with the product rather than living here: 12 of them in deploy/ops/prometheus-rules.yaml (plain Prometheus) and deploy/ops/prometheusrule.yaml (the Operator CRD), each with a runbook in deploy/ops/runbooks/. A test keeps every metric they reference in step with what the binary exports, keeps the two files identical, and fails if any alert loses its runbook.

This page used to carry its own examples with the same alert names and different thresholds, which is worse than having none: an operator who copied both got two rules called PrismCircuitBreakerOpen firing at different times. Use the shipped set.