Prometheus Metrics
PRISM exports metrics in Prometheus text exposition format on the admin API at GET /metrics.
Request Metrics
| Metric | Type | Description |
|---|---|---|
prism_requests_total{status="hit|miss|bypass|stale|error"} | counter | Total requests by cache status |
prism_inflight_requests | gauge | Currently processing requests |
prism_rate_limit_rejections_total | counter | Requests rejected by rate limiter (429) |
Render Metrics
| Metric | Type | Description |
|---|---|---|
prism_render_active | gauge | Currently active renders |
prism_render_queue_depth | gauge | Requests waiting for a Chrome tab |
prism_render_total | counter | Total renders completed |
prism_render_duration_ms | histogram | Render duration in milliseconds |
prism_shadow_renders_total | counter | Shadow renders completed (shadow mode) |
prism_content_validation_failures_total | counter | Renders that failed content validation |
prism_render_duration_ms histogram buckets: 100, 250, 500, 1000, 2500, 5000, 10000, 30000
Cache Metrics
| Metric | Type | Description |
|---|---|---|
prism_cache_entries | gauge | Number of entries in the render cache |
prism_cache_memory_bytes | gauge | Memory used by cached entries |
prism_cache_evictions_total | counter | Cache entries evicted (LRU or TTL) |
Chrome Tab Pool Metrics
| Metric | Type | Description |
|---|---|---|
prism_tab_pool_size | gauge | Total tabs in the pool |
prism_tab_pool_available | gauge | Idle tabs available for rendering |
prism_tab_crashes_total | counter | Tab crashes detected |
prism_tab_creates_total | counter | New tabs created |
prism_tab_recycles_total | counter | Tabs recycled after max renders |
prism_chrome_restarts_total | counter | Full Chrome process restarts |
Origin Metrics
| Metric | Type | Description |
|---|---|---|
prism_origin_requests_total | counter | Requests forwarded to the origin server |
prism_origin_duration_sum_ms | counter | Cumulative origin request duration (ms) |
Warmup Metrics
| Metric | Type | Description |
|---|---|---|
prism_warmup_urls_processed | counter | URLs successfully warmed up |
prism_warmup_urls_failed | counter | URLs that failed during warmup |
Circuit Breaker
| Metric | Type | Description |
|---|---|---|
prism_circuit_breaker_state | gauge | 0 = closed (healthy), 1 = open (tripped), 2 = half-open (probing) |
System Metrics
| Metric | Type | Description |
|---|---|---|
prism_uptime_seconds | gauge | Seconds since PRISM started |
process_resident_memory_bytes | gauge | Resident memory (RSS) of the PRISM process |
process_virtual_memory_bytes | gauge | Virtual memory of the PRISM process |
process_cpu_seconds_total | counter | Total CPU time consumed |
process_open_fds | gauge | Open file descriptors |
process_max_fds | gauge | Maximum file descriptors (ulimit) |
process_threads | gauge | Number of OS threads |
Renderer Internals
| Metric | Type | Description |
|---|---|---|
prism_render_queue_wait_ms | histogram | Time a render waited for a free tab. Distinct from render duration: this is the queue, that is the work. |
prism_render_retries_total | counter | Renders retried after a stalled navigation. One request costing two tabs, so a climb here means the stall is getting more frequent. |
prism_render_failures_total{class} | counter | Render failures by what they say about the shared browser. Only class="systemic" feeds the circuit breaker. |
prism_navigation_duration_sum_ms | counter | Sum of Chrome navigation time — the origin plus its JavaScript. |
prism_html_extract_duration_sum_ms | counter | Sum of DOM serialisation time. Rising against flat navigation means pages got larger, not slower. |
prism_oversize_rejections_total{check} | counter | Documents refused for size, labelled by which check caught them. |
prism_passthrough_not_html_total | counter | Requests proxied unchanged because the origin served non-HTML. Not a failure, but each one cost a Chrome navigation first. |
prism_requests_rejected_total{reason} | counter | Requests shed, by reason: queue_full, circuit_open, rate_limited, render_failed. |
prism_rate_limited_total{limiter} | counter | Rate-limit rejections attributed to the ip or domain budget. Which one is shedding is usually the whole question. |
Chrome Process
| Metric | Type | Description |
|---|---|---|
prism_chrome_rss_bytes | gauge | Resident memory of the browser process and all its descendants. process_resident_memory_bytes covers PRISM only, and Chrome is where growth happens. |
prism_chrome_rss_age_seconds | gauge | Age of the RSS sample above. Exported so a stale gauge reads as stale rather than as a current measurement. |
prism_chrome_recycles_total | counter | Proactive recycles on age or render budget — routine hygiene, unlike prism_chrome_restarts_total, which counts failures. |
Licence
| Metric | Type | Description |
|---|---|---|
prism_license_valid | gauge | 1 when the licence allows rendering. Stays 1 throughout the grace window. |
prism_license_grace_active | gauge | 1 exactly while the grace window is open — the interval where renewing still helps. The only metric that shows it. |
prism_license_info{status} | gauge | Licence state as a label: valid, trial, expired, no_license, invalid. |
Bots and Fleet
| Metric | Type | Description |
|---|---|---|
prism_bot_verification_total{outcome} | counter | FCrDNS checks of claimed crawlers: verified, spoofed, unverifiable. Rising spoofed is an attack becoming visible; rising unverifiable is your resolver degrading. |
prism_peer_purge_failures_total | counter | Purge fan-outs that failed to reach a replica. Until durable purge lands, the only sign a purge did not converge. |
prism_no_transform_ignored_total | counter | Responses rendered although the origin sent Cache-Control: no-transform. Counted whether or not cache.honor_no_transform is set, so ignoring the directive is a visible decision rather than an omission. Non-zero on a deployment that cares about the letter of RFC 9111 §5.2.2.6 is the signal to turn that setting on. |
prism_background_task_failures_total | counter | Supervisor tasks that exited unexpectedly. Any non-zero value means a capability (crash recovery, orphan cleanup, limiter sweeps) is silently gone. |
Scrape Configuration
Add PRISM to your Prometheus scrape_configs:
scrape_configs:
- job_name: prism
scrape_interval: 15s
static_configs:
- targets: ['127.0.0.1:4001']
# If bearer_token is configured:
authorization:
type: Bearer
credentials: your-secret-token
Grafana Dashboard
Import this JSON snippet as a Grafana dashboard to get started with key panels:
{
"title": "Trident PRISM",
"panels": [
{
"title": "Cache Hit Rate",
"type": "stat",
"targets": [
{
"expr": "rate(prism_requests_total{status=\"hit\"}[5m]) / (rate(prism_requests_total{status=\"hit\"}[5m]) + rate(prism_requests_total{status=\"miss\"}[5m]))",
"legendFormat": "Hit Rate"
}
],
"fieldConfig": {
"defaults": { "unit": "percentunit", "thresholds": { "steps": [
{ "value": 0, "color": "red" },
{ "value": 0.7, "color": "yellow" },
{ "value": 0.9, "color": "green" }
]}}
}
},
{
"title": "Render Duration (p50 / p95 / p99)",
"type": "timeseries",
"targets": [
{
"expr": "histogram_quantile(0.50, rate(prism_render_duration_ms_bucket[5m]))",
"legendFormat": "p50"
},
{
"expr": "histogram_quantile(0.95, rate(prism_render_duration_ms_bucket[5m]))",
"legendFormat": "p95"
},
{
"expr": "histogram_quantile(0.99, rate(prism_render_duration_ms_bucket[5m]))",
"legendFormat": "p99"
}
],
"fieldConfig": { "defaults": { "unit": "ms" } }
},
{
"title": "Active Renders & Queue Depth",
"type": "timeseries",
"targets": [
{ "expr": "prism_render_active", "legendFormat": "Active" },
{ "expr": "prism_render_queue_depth", "legendFormat": "Queued" }
]
},
{
"title": "Tab Pool",
"type": "timeseries",
"targets": [
{ "expr": "prism_tab_pool_size", "legendFormat": "Total" },
{ "expr": "prism_tab_pool_available", "legendFormat": "Available" }
]
},
{
"title": "Cache Entries & Memory",
"type": "timeseries",
"targets": [
{ "expr": "prism_cache_entries", "legendFormat": "Entries" },
{ "expr": "prism_cache_memory_bytes", "legendFormat": "Memory (bytes)" }
]
},
{
"title": "Circuit Breaker State",
"type": "stat",
"targets": [
{ "expr": "prism_circuit_breaker_state", "legendFormat": "State" }
],
"fieldConfig": {
"defaults": {
"mappings": [
{ "type": "value", "options": { "0": { "text": "CLOSED", "color": "green" } } },
{ "type": "value", "options": { "1": { "text": "OPEN", "color": "red" } } },
{ "type": "value", "options": { "2": { "text": "HALF-OPEN", "color": "yellow" } } }
]
}
}
},
{
"title": "Origin Latency (avg)",
"type": "timeseries",
"targets": [
{
"expr": "rate(prism_origin_duration_sum_ms[5m]) / rate(prism_origin_requests_total[5m])",
"legendFormat": "Avg Origin Latency (ms)"
}
],
"fieldConfig": { "defaults": { "unit": "ms" } }
},
{
"title": "Process Memory",
"type": "timeseries",
"targets": [
{ "expr": "process_resident_memory_bytes", "legendFormat": "RSS" },
{ "expr": "process_virtual_memory_bytes", "legendFormat": "Virtual" }
],
"fieldConfig": { "defaults": { "unit": "bytes" } }
}
]
}
Alerts
The alerting rules ship with the product rather than living here: 12 of them
in deploy/ops/prometheus-rules.yaml (plain Prometheus) and
deploy/ops/prometheusrule.yaml (the Operator CRD), each with a runbook in
deploy/ops/runbooks/. A test keeps every metric they reference in step with
what the binary exports, keeps the two files identical, and fails if any
alert loses its runbook.
This page used to carry its own examples with the same alert names and
different thresholds, which is worse than having none: an operator who copied
both got two rules called PrismCircuitBreakerOpen firing at different
times. Use the shipped set.