Bot Traffic · September 11, 2026

Is GPTBot Hitting Your Site 4,200 Times a Day? (Probably Yes)

If you're relying on Google Analytics to track AI crawler traffic, you're missing almost all of it. Here's what server logs actually show — and the numbers are bigger than most site owners expect.

By the Wrenda team · This article was generated with AI. Figures are sourced where cited below.

If you asked most site owners how much traffic GPTBot sends to their site, they'd probably say something like "a bit, maybe?" They've heard about blocking it, maybe added a disallow rule. But very few people have actually measured it.

Here's a number that tends to stop the room: 4,200. That's the median daily HTTP request count from GPTBot to a typical production website in 2026, measured across 12 real production sites over 30 days. One request roughly every 20 seconds, around the clock, every day — not a one-time crawl but steady background traffic.

The catch: if your analytics stack is client-side (Google Analytics, Mixpanel, PostHog, basically anything that runs in the browser), you're seeing none of it. AI crawlers don't execute JavaScript. They fetch raw HTML and move on. For most teams, this traffic is completely invisible in their dashboards — even as it consumes real bandwidth and real compute every day.

What do AI bots actually look like in your server logs?

Server log analysis is the only reliable way to see this traffic. When you look at it, the picture is more stratified than you'd expect.

Median Daily AI Bot Requests Per Site (2026)
Measured across 12 production sites over 30 days. Includes B2B SaaS, ecommerce, agency, and publisher properties.

GPTBot leads by a significant margin at 4,200 daily requests per site. ClaudeBot runs at roughly 1,800 requests per day, PerplexityBot at 980, and Google-Extended (Google's AI training crawler, separate from Googlebot) at around 540. These aren't isolated events — they're steady, daily baselines that continue regardless of whether you're tracking them.

Beyond volume, it matters a lot what these bots are actually trying to do. GPTBot and ClaudeBot are training crawlers — they're collecting data to improve AI models, not to answer user queries in real time. PerplexityBot and OAI-SearchBot are AI-search crawlers — they index your content to surface as live answers when someone asks a question. Training crawlers account for roughly 67.5% of all AI-driven traffic by volume. Search crawlers are smaller in absolute terms but much more directly tied to referral visits actually reaching your site.

Why are some sites blocking AI crawlers — and is it working?

The robots.txt blocking story has evolved a lot since 2023. Back then, only around 5% of the top 1,000 websites blocked GPTBot. That number climbed fast — and has now plateaued at roughly 25%.

GPTBot Blocking Rate Among Top 1,000 Sites: 2023 vs 2026
Share of the top 1,000 websites disallowing GPTBot in robots.txt. Growth plateaued after mid-2024.

The more interesting trend isn't the raw blocking rate, it's the split that's emerged. Around 30% of top sites now take a targeted approach: blocking training crawlers (GPTBot, ClaudeBot, CCBot, Google-Extended) while explicitly allowing AI-search crawlers (PerplexityBot, OAI-SearchBot). The reasoning is pretty clear-cut — training crawlers could use your content in datasets without ever citing your site, while search crawlers actually refer users back to you. Same content, very different outcomes.

One thing worth knowing: roughly 39.5% of GPTBot disallow rules in robots.txt aren't backed by any technical enforcement. The bot typically honours the policy, but site owners who care about this should enforce it at the server or network level, not just write it in a text file.

Does the type of site you run change the picture?

Significantly. B2B SaaS and technical documentation sites currently see the highest AI referral rates of any measured vertical — around 2.8% of total visits arriving from AI sources. That's more than 10 times the rate for typical e-commerce sites, which sit below 0.2% of sessions despite receiving high absolute crawler volume.

Why the gap? Content type. AI assistants tend to surface content that directly answers a specific question — pricing comparisons, API docs, how-to guides, feature breakdowns. Product listing pages and category pages get crawled heavily but cited much less often in actual AI responses.

What's actually worth doing about this?

If you haven't looked at your server logs with AI bot traffic isolated, that's the obvious starting point. It doesn't require special tooling — filter requests by user-agent string (GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Google-Extended) and you'll see exactly what's hitting your site and which pages are getting the most attention. Most teams are genuinely surprised by the absolute numbers the first time they see them.

Once you know the volume, the decisions become fairly clear. If you want AI assistants to recommend your site when users search for what you offer, make sure AI-search crawlers can access your key content cleanly and quickly. Blocked pages and slow-loading HTML both hurt your chances of being cited.

If you're concerned about your content being used for model training without attribution, targeted robots.txt disallows for training crawlers are effective — but make sure there's actual enforcement behind them, not just the policy.

What doesn't make sense for most sites is a blanket block on all AI crawlers. Training crawlers and search crawlers have very different implications for your discoverability, and treating them the same means you're probably blocking traffic that would have helped you. The question isn't whether AI bots are crawling your site right now — they almost certainly are. The question is whether your access policy actually reflects what you want to happen.

Sources

  1. Agentic Crawler Behavior: 30-Day Site Log Study
  2. AI Crawler Volume Growth 2022-2026: GPTBot +305% YoY
  3. The AI Crawler Block Index
  4. 2026 State of AI Traffic & Cyberthreat Benchmark Report (HUMAN Security)
  5. AI Referral Traffic by Industry: 2026 Data (Similarweb)