1001 SEO Media
All posts
By 1001 SEO MediaGEOAI searchAI crawlersmeasurement

AI Crawler Activity in Website Logs: Cloudflare AI Bots, Visibility Tools, Google AI Overviews Crawler, and Bing Copilot Crawler Analytics (2026)

Which bot really feeds AI Overviews, Copilot and ChatGPT, what Cloudflare shows you natively, and how we tie crawler logs to prompts and visits on client programs.

The first question on most crawler audits we run is some version of "where is the AI Overviews bot in our logs?" There isn't one. Before comparing any tools, it pays to get the map right, because a fair share of the crawler reports we inherit filter for user agents that will never appear.

Which crawler feeds which AI surface

Google AI Overviews and AI Mode. Google's common crawlers documentation (last updated July 2026) says Googlebot's crawl preferences apply to Google Search, "including Discover and all Google Search features." AI Overviews and AI Mode are Search features. The crawler that feeds them is plain Googlebot. Block Googlebot to keep a page out of AI Overviews and you have also pulled it from the regular results.

Google-Extended trips people up. It is a robots.txt token that controls whether content is used for Gemini training and for grounding in Gemini Apps and Vertex AI. Google says it "doesn't have a separate HTTP request user agent string" and that it does not affect inclusion in Google Search or act as a ranking signal. You will never find Google-Extended in a log file. If a dashboard reports Google-Extended hits, it is counting something else.

GoogleOther is a generic Google crawler that, per the same documentation, doesn't affect any specific product. It's real and it shows up. It isn't an AI Overviews bot either.

Bing Copilot. Microsoft's webmaster guidelines say Bing and Copilot search experiences "rely on the same core crawling, indexing, and ranking foundation as traditional search." So the Copilot crawler in your logs is Bingbot, and the guidelines list blocking Bingbot in robots.txt as a mistake. If a client wants a page crawled but kept out of Copilot answers, the levers are meta directives. NOARCHIVE prevents content from being used in Copilot responses and grounding results. NOCACHE limits Copilot to the URL, title and snippet. NOINDEX keeps the URL out of Bing, Copilot and grounding results altogether.

ChatGPT. OpenAI documents GPTBot and OAI-SearchBot as separate robots, one for training and one for search, and ChatGPT-User covers fetches a person triggers inside ChatGPT. Our OAI-SearchBot notes cover what that means for robots.txt.

One more thing before counting anything: user agents can be spoofed. Verify Googlebot by reverse DNS against googlebot.com or against Google's published IP ranges. Verify Bingbot with Microsoft's Verify Bingbot tool or reverse DNS. A scraper that calls itself Googlebot is a classic source of "Google tripled its crawl overnight" alarms.

What Cloudflare shows without any extra tooling

If the site runs behind Cloudflare, start with AI Crawl Control, which used to be called AI Audit. It is available on all plans. The Overview, Crawlers and Metrics tabs break requests down by crawler and by operator (OpenAI, Microsoft, Google, ByteDance, Anthropic, Meta). You see allowed versus unsuccessful requests, robots.txt violations, status codes, paths and hosts, and you can allow or block each crawler. Pay per crawl is in private beta.

The plan tier changes what you get. On Free, detection works from the user-agent string only and the Metrics tab covers the past 24 hours. Referral counts are limited to paid plans. Enterprise with Bot Management adds detection IDs, configurable timeframes and pay per crawl. Engineering can also pull the same data through Cloudflare's GraphQL Analytics API.

For a quick check like "is GPTBot getting 403s on the pricing page," that is enough. We open it in week one of every program where the client is on Cloudflare.

Where a crawler log stops being useful

A log answers one question well: did the bot fetch the URL, and what status did it get back. It doesn't say whether that fetch ended up in an answer, for which prompt, or whether anyone clicked through and converted. Cloudflare doesn't try to connect a crawl to a prompt, a citation or a visit. That's fine. It's a CDN.

The problem comes when teams read crawler volume as a visibility score. More PerplexityBot hits feel like progress. Sometimes they are. Sometimes the bot is refetching a 404 that a stale sitemap still lists, and the count is going up because something is broken.

How we wire this up for clients

We route crawler logs into Promptwatch, the platform we run client programs on, because its Agent Analytics lives in the same project as the prompts and the citations. Its documented crawler coverage is ChatGPTBot, ClaudeBot, PerplexityBot, GoogleOther and the Meta AI crawler, with error tracking per URL and a crawl-to-citation path. In practice that path reads as: this page was fetched on this date, then cited in the answer to this prompt. Visitor analytics, installed as a small script or a GTM template, covers the last step with AI-referred sessions and conversions.

Cloudflare connects through Logpush on Enterprise or an auto-deployed Worker on other plans. The DNS record has to be orange-cloud proxied or there is nothing to ship. We wrote up the steps in connecting Cloudflare crawler logs. AWS CloudFront, Fastly, Vercel, Netlify, Akamai, Google Cloud CDN and a custom HTTP endpoint are supported too.

For Google and Bing we keep the two questions apart on purpose. Googlebot and Bingbot health gets read in the raw CDN or server logs. Presence in AI Overviews, AI Mode and Copilot gets read on the answer side, since Promptwatch monitors those surfaces next to ChatGPT, Gemini, Claude and Perplexity. We don't tell clients there is a Google AI crawler in that list, because Google doesn't run one for AI Overviews.

A plan note we put in every proposal. Crawler logs start at Professional ($245/mo, 25M crawler logs). Business is $579/mo with 100M. Agency plans include them from Kick-off ($199/mo, 10M). Essential ($95/mo) has visitor analytics but no listed crawler-log allowance, so we don't sell a crawler audit on it.

The checklist we run in week one

  1. Pull a week of logs and verify Googlebot and Bingbot by reverse DNS before counting a single hit.
  2. Confirm robots.txt allows Googlebot, Bingbot and OAI-SearchBot on pages the client wants cited. Treat GPTBot and Google-Extended as separate decisions, since both are about training.
  3. Open Cloudflare AI Crawl Control and sort unsuccessful requests by operator.
  4. Fix 4xx and 5xx responses on citation candidates before anyone rewrites copy.
  5. Connect the CDN to Promptwatch and watch which fetched pages become citations, and which citations send visitors.

If you'd like us to run that audit on your logs, email hello@1001seomedia.com with your CDN and plan tier.