OpenAI and Perplexity user-agent patterns on indie marketing sites

Recognize fetchers in logs, separate them from Googlebot, and avoid blocking paying customers by mistake.

All guides on this topic: AI & search crawlers on indie sites

Answer-engine crawlers change User-Agent strings without asking your permission. You are not expected to memorize every token, you are expected to group them into “AI answers” in analytics and not merge them with human Chrome.

Spoke in AI crawler guides. Start at the Main guide.

Why User-Agent classification is brittle

Vendors rotate strings, use generic compatible tokens, or fetch via third-party infra. Your classifier should update quarterly, not set-and-forget regex from a blog post in 2023.

OpenAI-family fetchers

Often include GPTBot, ChatGPT-User, or OAI-Search style tokens depending on product surface. Policy pages describe allow/disallow; analytics describes volume.

Perplexity fetchers

Distinct tokens from search indexers. Low counts on tiny sites are normal, Perplexity indexes selectively.

Do not block human preview tools

Some “fetch” UAs power link previews in chat apps. Blocking may break sharing cards. Prefer path-level rules on /docs only if needed.

Compare to Googlebot

Googlebot volume usually dwarfs OpenAI on established docs. If OpenAI exceeds Google on /pricing, something is scraping aggressively, investigate IP rate limits.

Log sampling for verification

Once a month, export raw User-Agent for top 20 bot hits. Update classifier rules in visitor-classify (your codebase), not in Cloudflare alone, analytics and firewall should agree.

CDN vs application analytics

Cloudflare “bot score” ≠ marketing analytics taxonomy. Use one source of truth for founder-facing charts.

Privacy framing

Bot hits are not GDPR personal data in most setups; still avoid logging full IP in public exports.

When counts jump after press

Tech press mentions your API docs, answer bots and humans both rise. Pin the week on marketing pulse guides.

False human classification

Rare headless clients look like bots. Do not auto-block without revenue impact review.

Worked week

Docs site: Googlebot 12k, OpenAI 340, Perplexity 22, humans 1.8k. Action: improve human CTA on top doc page, not robots.txt war.

Hub

Cluster

Names change; categories stay. Track categories for decisions, strings for engineering tickets.

Maintainer mindset for classifier rules

Treat User-Agent rules like dependency updates: quarterly review, changelog in repo, test against last month’s raw log sample. Vendors publish allowlists; your analytics taxonomy should map strings → category in one module so marketing copy does not embed stale regex.

OpenAI surfaces (conceptual, not legal)

Fetch traffic may come from training crawlers, browsing tools, or search-augmented products with different policy pages. Analytics buckets collapse them into AI answers for founder sanity. Engineering tickets can list exact strings when blocking or allowing paths.

Perplexity at indie scale

Double-digit daily hits on a 3k-visitor site is normal. Zero for months then a burst after a niche blog post is also normal. Compare to AI crawler guide sidebar rankings, Perplexity is often fourth place, not first.

Path-level bot tables

PathGooglebotOpenAI-classHumans
/docs/apihighmediummedium
/pricinglowlowhigh
/blog/launchmediumlowhigh

If OpenAI-class is high on /pricing, verify it is not a misclassified headless checkout test before blocking.

Blocking checklist (do not rush)

  1. Confirm misclassification with raw logs.
  2. Check revenue impact (none for bots).
  3. Prefer robots.txt path rules over firewall country blocks.
  4. Re-measure tab counts 7d after change.
  5. Document decision in internal wiki.

Link previews vs crawlers

Chat apps fetch Open Graph tags with dissimilar User-Agents. Blocking “unknown fetchers” can break preview cards in Slack, sales team notices before analytics does. Test share card after any WAF rule change.

Correlation with human signup paths

Spike in OpenAI-class fetches on /docs/quickstart plus rise in human /docs → /pricing paths suggests answer engines surface your quickstart. Add ethical CTA on that page; see docs before checkout.

Enterprise log tools vs indie analytics

Datadog may show millions of bot requests; marketing chart shows thousands of classified pageviews, sampling and dedupe differ. Pick one founder-facing source (your product analytics) for weekly reviews.

Incident vs trend

One day 500% spike: deploy or press. Seven days gradual climb: indexing interest. Four weeks flat high Training bucket: investigate scrape per training noise guide.

Update playbook when strings change

  1. Export top 50 unknown User-Agents.
  2. Search vendor changelog / status pages.
  3. Add mapping or explicit unknown bucket.
  4. Deploy classifier.
  5. Note in team channel, not every string deserves a blog post.

Related topics

Closing honesty

You cannot opt out of the entire answer-engine ecosystem with regex alone. You can measure it, rate-limit abuse, and keep human metrics primary. Categories in charts exist so you spend Friday on landing page RPV, not on bot conspiracy threads.

Historical string archive

Keep a docs/bot-ua-changelog.md in repo with date, string observed, category assigned. When a vendor rotates UA, you diff git blame instead of guessing when charts broke.

Headless Chrome false positives

Puppeteer-based monitoring on your own site can inflate “bot” counts if classified as unknown. Exclude your monitoring IPs or tag synthetic checks utm_source=monitoring on human paths only.

Perplexity vs Google user intent

Googlebot-heavy weeks often precede organic human clicks. Perplexity-heavy weeks may precede direct traffic from people who read answers elsewhere, harder to attribute. Watch branded search and direct RPV together in revenue per visitor guides.

Docs paywall experiments

If you gate docs, AI fetchers may drop before humans complain. Measure both tabs and documentation checkout paths for two weeks minimum before declaring win.

Vendor status pages

Subscribe to crawler status RSS where offered. Outage explains flat lines better than SEO panic.

Community reports

Indie founders share UA strings in forums, verify against your logs before adding rules. Copy-paste errors propagate bad regex.

Rate limit tuning

When OpenAI-class fetchers hammer one IP range, rate limit at edge without global bot fight. Log blocked count separately so analytics and firewall stories align.

Annual review agenda

Q1: export top UA strings. Q2: update classifier. Q3: robots.txt policy check. Q4: ignore Training unless incident. Repeat.

Handoff to contractors

Give contractors category names, not raw regex files, unless they own analytics ingest. Miscommunication blocks Googlebot while “fixing GPTBot.”

Final worked decision

Indexing +12% WoW, OpenAI +8%, humans +3%, RPV flat → keep docs public, add CTA on top fetched page, no firewall changes. Write that sentence in your journal; it is a complete AI UA review.

When to revisit this guide

Re-read after any major docs redesign, robots.txt edit, or press hit mentioning your API. Otherwise quarterly classifier maintenance is enough, your energy belongs on UTM discipline and checkout metadata, not daily UA archaeology.

Sidebar ranking as a health check

When OpenAI-class counts exceed Google on a docs property older than six months, treat it as a classifier audit trigger, not as “we beat Google at SEO.” Compare raw logs, fix mappings, then re-check the AI crawler guide 24h chart. More guides: AI crawler visibility. When in doubt, classify conservatively into AI answers and fix false positives in logs, false negatives in public charts are easier to explain than blocking real users.

More on this topic