Answer-engine crawlers change User-Agent strings without asking your permission. You are not expected to memorize every token, you are expected to group them into “AI answers” in analytics and not merge them with human Chrome.
Spoke in AI crawler guides. Start at the Main guide.
Why User-Agent classification is brittle
Vendors rotate strings, use generic compatible tokens, or fetch via third-party infra. Your classifier should update quarterly, not set-and-forget regex from a blog post in 2023.
OpenAI-family fetchers
Often include GPTBot, ChatGPT-User, or OAI-Search style tokens depending on product surface. Policy pages describe allow/disallow; analytics describes volume.
Perplexity fetchers
Distinct tokens from search indexers. Low counts on tiny sites are normal, Perplexity indexes selectively.
Do not block human preview tools
Some “fetch” UAs power link previews in chat apps. Blocking may break sharing cards. Prefer path-level rules on /docs only if needed.
Compare to Googlebot
Googlebot volume usually dwarfs OpenAI on established docs. If OpenAI exceeds Google on /pricing, something is scraping aggressively, investigate IP rate limits.
Log sampling for verification
Once a month, export raw User-Agent for top 20 bot hits. Update classifier rules in visitor-classify (your codebase), not in Cloudflare alone, analytics and firewall should agree.
CDN vs application analytics
Cloudflare “bot score” ≠ marketing analytics taxonomy. Use one source of truth for founder-facing charts.
Privacy framing
Bot hits are not GDPR personal data in most setups; still avoid logging full IP in public exports.
When counts jump after press
Tech press mentions your API docs, answer bots and humans both rise. Pin the week on marketing pulse guides.
False human classification
Rare headless clients look like bots. Do not auto-block without revenue impact review.
Worked week
Docs site: Googlebot 12k, OpenAI 340, Perplexity 22, humans 1.8k. Action: improve human CTA on top doc page, not robots.txt war.
Hub
Names change; categories stay. Track categories for decisions, strings for engineering tickets.
Maintainer mindset for classifier rules
Treat User-Agent rules like dependency updates: quarterly review, changelog in repo, test against last month’s raw log sample. Vendors publish allowlists; your analytics taxonomy should map strings → category in one module so marketing copy does not embed stale regex.
OpenAI surfaces (conceptual, not legal)
Fetch traffic may come from training crawlers, browsing tools, or search-augmented products with different policy pages. Analytics buckets collapse them into AI answers for founder sanity. Engineering tickets can list exact strings when blocking or allowing paths.
Perplexity at indie scale
Double-digit daily hits on a 3k-visitor site is normal. Zero for months then a burst after a niche blog post is also normal. Compare to AI crawler guide sidebar rankings, Perplexity is often fourth place, not first.
Path-level bot tables
| Path | Googlebot | OpenAI-class | Humans |
|---|---|---|---|
/docs/api | high | medium | medium |
/pricing | low | low | high |
/blog/launch | medium | low | high |
If OpenAI-class is high on /pricing, verify it is not a misclassified headless checkout test before blocking.
Blocking checklist (do not rush)
- Confirm misclassification with raw logs.
- Check revenue impact (none for bots).
- Prefer
robots.txtpath rules over firewall country blocks. - Re-measure tab counts 7d after change.
- Document decision in internal wiki.
Link previews vs crawlers
Chat apps fetch Open Graph tags with dissimilar User-Agents. Blocking “unknown fetchers” can break preview cards in Slack, sales team notices before analytics does. Test share card after any WAF rule change.
Correlation with human signup paths
Spike in OpenAI-class fetches on /docs/quickstart plus rise in human /docs → /pricing paths suggests answer engines surface your quickstart. Add ethical CTA on that page; see docs before checkout.
Enterprise log tools vs indie analytics
Datadog may show millions of bot requests; marketing chart shows thousands of classified pageviews, sampling and dedupe differ. Pick one founder-facing source (your product analytics) for weekly reviews.
Incident vs trend
One day 500% spike: deploy or press. Seven days gradual climb: indexing interest. Four weeks flat high Training bucket: investigate scrape per training noise guide.
Update playbook when strings change
- Export top 50 unknown User-Agents.
- Search vendor changelog / status pages.
- Add mapping or explicit
unknownbucket. - Deploy classifier.
- Note in team channel, not every string deserves a blog post.
Related topics
- Googlebot sitemap bursts
- Path funnels hub for human outcomes
- UTM guides when humans arrive from tagged campaigns after reading AI-surfaced docs
Closing honesty
You cannot opt out of the entire answer-engine ecosystem with regex alone. You can measure it, rate-limit abuse, and keep human metrics primary. Categories in charts exist so you spend Friday on landing page RPV, not on bot conspiracy threads.
Historical string archive
Keep a docs/bot-ua-changelog.md in repo with date, string observed, category assigned. When a vendor rotates UA, you diff git blame instead of guessing when charts broke.
Headless Chrome false positives
Puppeteer-based monitoring on your own site can inflate “bot” counts if classified as unknown. Exclude your monitoring IPs or tag synthetic checks utm_source=monitoring on human paths only.
Perplexity vs Google user intent
Googlebot-heavy weeks often precede organic human clicks. Perplexity-heavy weeks may precede direct traffic from people who read answers elsewhere, harder to attribute. Watch branded search and direct RPV together in revenue per visitor guides.
Docs paywall experiments
If you gate docs, AI fetchers may drop before humans complain. Measure both tabs and documentation checkout paths for two weeks minimum before declaring win.
Vendor status pages
Subscribe to crawler status RSS where offered. Outage explains flat lines better than SEO panic.
Community reports
Indie founders share UA strings in forums, verify against your logs before adding rules. Copy-paste errors propagate bad regex.
Rate limit tuning
When OpenAI-class fetchers hammer one IP range, rate limit at edge without global bot fight. Log blocked count separately so analytics and firewall stories align.
Annual review agenda
Q1: export top UA strings. Q2: update classifier. Q3: robots.txt policy check. Q4: ignore Training unless incident. Repeat.
Handoff to contractors
Give contractors category names, not raw regex files, unless they own analytics ingest. Miscommunication blocks Googlebot while “fixing GPTBot.”
Final worked decision
Indexing +12% WoW, OpenAI +8%, humans +3%, RPV flat → keep docs public, add CTA on top fetched page, no firewall changes. Write that sentence in your journal; it is a complete AI UA review.
When to revisit this guide
Re-read after any major docs redesign, robots.txt edit, or press hit mentioning your API. Otherwise quarterly classifier maintenance is enough, your energy belongs on UTM discipline and checkout metadata, not daily UA archaeology.
Sidebar ranking as a health check
When OpenAI-class counts exceed Google on a docs property older than six months, treat it as a classifier audit trigger, not as “we beat Google at SEO.” Compare raw logs, fix mappings, then re-check the AI crawler guide 24h chart. More guides: AI crawler visibility. When in doubt, classify conservatively into AI answers and fix false positives in logs, false negatives in public charts are easier to explain than blocking real users.