Are AI Crawlers Taking More Than They Give? The Numbers Behind the Gap
AI crawlers now account for 22% of all verified bot traffic — but ClaudeBot returns one referral for every 1,917 pages it crawls. So what exactly are they giving back?
A ratio that should make you stop and think
ClaudeBot crawled 1,917 of your pages last month for every single referral visit it sent back. GPTBot: 217 to 1. PerplexityBot: roughly 111 to 1. Traditional search — the channel you have been optimising for the last decade — returns a visit for about every 4.9 pages it indexes.
So what exactly are you getting from all this AI crawler activity?
That is the uncomfortable question at the heart of the crawl-to-referral gap: the growing distance between how much AI bots consume from your site and how little traffic they return. The gap has narrowed over the past eight months, but it is still enormous. And if you do not understand which bots are responsible for which part of that equation, you are flying blind on a decision that affects your content strategy, your infrastructure costs, and your visibility in AI search results.
Where the numbers come from
The crawl-to-referral ratios in this post come from a GEO data study tracking over 2,000 websites through Q2 and Q3 2026, measuring pages crawled per AI bot user-agent against referral sessions those bots sent back. Traffic composition figures come from CDN infrastructure reports covering verified bot categories through May and June 2026. Publisher blocking rates are drawn from a robots.txt compliance analysis across the open web and publisher segments.
How big a share of your traffic is actually AI now?
Two years ago, AI crawlers made up roughly 10% of all verified bot traffic. By May 2026, that share hit 22.3%. Add AI search bots — which operate separately from training crawlers — and AI-related activity accounts for approximately 26.7% of verified bot traffic. That is more than one in four bot requests coming from an AI system of some kind.
Zoom out to all web traffic and it gets more striking. Automated requests now represent 57.5% of all HTML web traffic versus 42.5% from human browsers. The bots have collectively overtaken humans in share of raw requests. And AI crawlers are the fastest-growing slice of that automated traffic — real-time AI agent requests, the kind where a bot acts on behalf of a user in real time, grew an estimated 7,800% year over year, albeit from a small base.
This is not a trend to check once and move on from. The AI crawler landscape is changing month to month, with new user-agents appearing, old ones scaling back, and the balance between training and search traffic shifting as AI products evolve.
So what does 1,917-to-1 actually mean for a content site?
The crawl-to-referral ratio is simple to define but the implications take a minute to sink in. It measures how many of your pages a bot indexes before its parent platform sends a single user back to your site. A ratio of 100:1 means the bot visited 100 pages and the platform directed one visitor back to you.
Here is the August 2026 picture across the major crawlers:
- Traditional search: approximately 4.9:1
- PerplexityBot: approximately 111:1
- GPTBot (training crawler): 217:1
- ClaudeBot: 1,917:1
PerplexityBot is already 23 times worse than traditional search. ClaudeBot is nearly 400 times worse. And this is after significant improvement — ClaudeBot was sitting closer to 23,000:1 in early Q1 2026 before its associated search product began generating meaningful referral traffic.
The improvement trend is real and worth noting. When AI platforms ship products that actively surface cited links to users, their crawl-to-referral ratios drop dramatically. The question is whether your site is positioned to appear in those citations when they do.
Why training bots and search bots are a different problem entirely
Here is the split that most site owners miss: 51.8% of AI crawler requests in May 2026 were for training purposes, while only 9.3% were attributable to real-time search. These are fundamentally different systems with different user-agents and different return policies.
Training crawlers consume your content to improve an underlying model. They send almost no referrals back to you. Search crawlers index your content for real-time retrieval and are the ones generating nearly all the AI-attributed referral traffic that occasionally shows up in your analytics.
Most sites treat both identically in robots.txt, either allowing all or blocking all. But if your goal is to appear in AI search results while protecting your content from training pipelines, you need to distinguish between a training user-agent and its search-oriented counterpart. That requires reading the documentation for each major AI platform and writing separate rules accordingly. Most site configurations do not do this today.
Publishers are pushing back — but the barrier is porous
Publishers are not standing still. In Q2 2025, 3.63% of AI bot requests hit a 403 Forbidden response. By Q2 2026, that figure was 8.56% — more than double in twelve months. Among news publishers specifically, 82.4% of sites with a robots.txt file restrict at least one AI crawler.
But here is where it gets complicated. Of 592 sites that explicitly disallow GPTBot in their robots.txt, around 234 — roughly 39.5% — still served GPTBot a live 200 response when it made a request. The disallow directive is a convention, not a technical barrier. Nearly four in ten sites that have said no are not actually enforcing it at the server level.
This matters because robots.txt-only blocks give you a false sense of security. If you want to actually control what AI crawlers receive from your site, you need enforcement at the infrastructure layer — not just a file.
What should site owners actually do about this?
The first thing worth acknowledging is that a significant chunk of your AI-driven traffic is already invisible in your analytics. Research covering 446,000 sessions found that 70.6% of AI-driven visits arrive without a referrer header, which means your analytics tool classifies them as direct traffic. If your direct channel has grown unexpectedly over the past year, some of that growth is AI referrals that your tools cannot identify.
When AI traffic is attributable, it converts at a notable premium. Sign-up conversion rates for identified AI referrals average around 1.66%, compared to roughly 0.15% for organic search — an 11-times difference. The traffic quality, when you can see it, is high. The challenge is identifying enough of it to act on.
The practical implication: decisions about which AI crawlers to allow, block, or optimise for need to be specific. Blanket blocks protect your training data but cost you AI search visibility. Blanket allowances maximise your AI search surface but expose your content to training pipelines that give you little in return.
The crawl-to-referral ratios in this post are a starting point. When the math is 1,917:1, knowing which bot is responsible — and whether a rule change would actually move the referral number — is the difference between a deliberate strategy and an expensive accident. The exchange rate is improving, but it is still far enough from balanced that understanding it should be a standing part of how you manage your site.
Sources
- GEO Data Report 2026: Which AI Crawlers & LLM Bots Take the Most and Give the Least?
- Crawl-to-Refer Ratio 2026: AI Bots Take 4,580, Give 1
- AI agents now make up the majority of web traffic: What developers need to change
- Do News Sites Block AI Crawlers? robots.txt Data
- AI Search Traffic Report 2026: ChatGPT Referral Share Trends