Skip to main content

Bot Detection Configuration

The [bot] section configures how PRISM identifies bot user agents. In bot-only mode, only requests from detected bots are rendered; all other requests are proxied directly to the origin.

TOML Example​

[bot]
patterns = [
# Search engines
"Googlebot",
"Googlebot-Image",
"Googlebot-Video",
"Googlebot-News",
"Storebot-Google",
"Google-InspectionTool",
"GoogleOther",
"bingbot",
"Baiduspider",
"YandexBot",
"DuckDuckBot",
"DuckAssistBot",
"Slurp",
"Applebot",
"Applebot-Extended",
"PetalBot",
"Sogou",
"SeznamBot",
"Amazonbot",
"Bravebot",
# AI/LLM crawlers
"GPTBot",
"ChatGPT-User",
"OAI-SearchBot",
"ClaudeBot",
"Claude-User",
"Claude-SearchBot",
"anthropic-ai",
"PerplexityBot",
"Bytespider",
"meta-externalagent",
"Meta-ExternalFetcher",
"FacebookBot",
"CCBot",
"DeepSeekBot",
"cohere-ai",
"Diffbot",
"YouBot",
"PhindBot",
"FirecrawlAgent",
"Timpibot",
"ImagesiftBot",
# Social / link preview
"facebookexternalhit",
"Twitterbot",
"LinkedInBot",
"Pinterestbot",
"Discordbot",
"WhatsApp",
"TelegramBot",
"Slackbot",
"redditbot",
"Snap URL Preview",
"Bluesky",
"Mastodon",
"Viber",
"kakaotalk-scrap",
"Iframely",
"FlipboardProxy",
# SEO tools
"AhrefsBot",
"SemrushBot",
"MJ12bot",
"DotBot",
"DataForSeoBot",
"ContentKingApp",
"Screaming Frog",
"Embedly",
"Quora Link Preview",
# Archive
"ia_archiver",
]

Parameters​

ParameterTypeDefaultDescription
patternsArray of Strings(60+ patterns, see above)User-agent substrings that identify bot traffic

Detection Strategy​

PRISM uses a two-layer bot detection approach:

  1. Primary: isbot crate -- PRISM first checks the User-Agent header against the isbot library, which maintains a comprehensive, regularly updated database of known bot signatures.

  2. Fallback: patterns list -- If isbot does not match, PRISM checks whether the User-Agent contains any of the configured pattern strings as a substring. Matching is case-insensitive.

This dual approach ensures reliable detection even for new or niche crawlers not yet in the isbot database.

Detailed Explanation​

Pattern matching​

Each entry in the patterns array is treated as a case-insensitive substring match against the full User-Agent header. For example, the pattern "Googlebot" matches:

  • Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
  • Googlebot-Image/1.0
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1)

Default pattern categories​

The default list covers four major categories:

  • Search engines (20 patterns): Google, Bing, Baidu, Yandex, DuckDuckGo, Apple, and regional search engines
  • AI/LLM crawlers (21 patterns): GPTBot, ClaudeBot, PerplexityBot, DeepSeekBot, and other AI training/search crawlers
  • Social/link preview (16 patterns): Facebook, Twitter/X, LinkedIn, Discord, WhatsApp, Telegram, Slack, Reddit, Bluesky, Mastodon
  • SEO tools and archives (10 patterns): Ahrefs, Semrush, Screaming Frog, Internet Archive

Overriding the default list​

Setting patterns in your config replaces the entire default list. If you only want to add a custom bot, you must include the defaults plus your additions:

[bot]
patterns = [
"Googlebot",
"bingbot",
# ... include defaults you need ...
"MyCustomCrawler",
]

Example Use Cases​

Minimal bot list for testing​

[bot]
patterns = ["Googlebot", "bingbot"]

Adding a custom internal crawler​

[bot]
patterns = [
# Keep all defaults plus your custom bot
"Googlebot",
"bingbot",
"Baiduspider",
"YandexBot",
# ... other defaults ...
"InternalMonitorBot",
"MyCompanyCrawler",
]

Rendering for all traffic (no bot detection needed)​

If you use render-all mode, the bot patterns list is not consulted for routing decisions, but it is still used for analytics and the X-Prism-Bot response header.

[server]
mode = "render-all"

# Bot patterns are still used for tagging, not routing
[bot]
patterns = ["Googlebot", "bingbot"]

Verifying claimed crawlers (FCrDNS)​

patterns answers "does this User-Agent claim to be a bot"; nothing above answers "is the claim true". A spoofed Googlebot UA buys the expensive path — a render, a Chrome tab, a cache entry — for the cost of one header.

[bot.verify] checks claimed identities the way their operators document: the client IP's PTR record must land in the operator's domain (crawl-66-249-66-1.googlebot.com), and the forward lookup of that name must contain the same IP. Verdicts are cached per IP (verified 24h, spoofed 60m by default), so a real crawl costs one pair of lookups, not one per request.

[bot.verify]
enabled = true # observe: verdicts in metrics and logs
enforce = false # act: proven-spoofed crawler UAs are treated as humans

In bot-only mode, enforcement means a proven-spoofed request is proxied — never rendered, never occupying a Chrome tab. In render-all mode it is still rendered (as any human is) but keyed and logged as a human. Logging: spoofed is a warning, unverifiable an info line, verified debug-level; the metric carries all three unconditionally.

Three verdicts, deliberately asymmetric:

  • verified — the round trip completed and confirmed the claim.
  • spoofed — a completed check positively failed: PTR outside the operator's domain, no PTR at all, or the forward lookup did not contain the client IP. Only this verdict is ever enforced against.
  • unverifiable — DNS timed out or failed. Never enforced against: a degraded resolver must not demote real Googlebot traffic to the raw JavaScript shell.

Built-in rules cover Googlebot, bingbot, Applebot, Yandex and Baiduspider with each operator's published rDNS domains. An explicit rules key in config replaces that list rather than extending it — restate the defaults you want to keep. A user agent matching no rule — an SEO crawler, curl, your monitoring — claims nothing verifiable and is left exactly as patterns classified it.

Watch prism_bot_verification_total{outcome=...}: spoofed rising is an attack becoming visible; unverifiable rising is your resolver degrading — different alerts.