Why Are AI Bots Crawling Your Site 286,000 Times Without Sending You a Single Visitor?
AI crawlers now make up 22% of all bot traffic — and one major crawler was fetching 286,000 pages for every single referral it sent back. Here's what the real numbers look like, and what you should do about it.
At some point in January 2025, a major AI assistant's crawler fetched 286,000 pages from a typical publisher before sending back a single referral visitor. Not 286,000 pages across the whole web — 286,000 pages from one site, before one human clicked through. The crawl-to-referral ratio at that moment: 286,000 to one. By July 2026, that same bot had improved to 2,237:1. Better — but is that better enough to justify the server load? And should it change how you treat these crawlers?
Where does this data come from?
This breakdown draws on Fastly's threat research analysis of AI bot traffic patterns across billions of requests, HUMAN Security's 2026 AI traffic benchmark report, and Digital Applied's Q1-Q2 2026 bot traffic reference dataset. Crawl-to-referral ratios come from SEOmator's GEO data report, which tracks AI crawler volume against actual platform referrals measured at the domain level.
So who's actually hitting your site right now?
Googlebot still leads at 27.26% of verified AI-adjacent bot HTTP requests in May 2026 — no surprise there. Behind it: GPTBot at 11.48% and ClaudeBot at 9.73%, with a long tail of Bytespider, PerplexityBot, and others filling the rest.
Why does that distribution matter? GPTBot grew 305% year-over-year between May 2024 and May 2025. ClaudeBot's share climbed from around 6% to nearly 10% over the same window. By Q1 2026, AI crawlers as a category had reached 22% of all bot traffic — up from negligible levels just two years earlier. More than one in five bot-tier requests now comes from something that isn't a traditional search crawler.
The growth isn't plateauing. HUMAN Security's 2026 benchmark projects AI bots on track to surpass human traffic in total request volume before end of 2027.
Are these bots learning, or actually answering questions?
This is probably the most useful distinction that standard analytics setups completely miss.
Not every AI bot visiting your site is doing the same job. Training crawlers — bots collecting data to improve AI models — made up roughly 90% of all AI-driven traffic in January 2025. Real-time scrapers, which fetch content in response to live user queries, accounted for the remaining 10%. Agentic bots (AI agents browsing autonomously on behalf of users) barely registered.
By December 2025, Fastly's research and HUMAN Security's data tell consistent stories: training crawlers had fallen to around 74% of AI bot traffic. Real-time scrapers climbed to 24%. The newly emerged agentic category appeared at roughly 2%.
The practical implication of this split: a training crawler hitting your pages today might influence how an AI system describes your content for years. A real-time scraper is fetching your page right now because someone typed a question — and that platform might send that user to you, or might answer them directly without a click-through. Two very different interactions that call for different responses.
What's behind the GPTBot and ClaudeBot share swap?
If you've watched your logs this year, you've probably noticed these two bots trading places. In April 2026, ClaudeBot led at 11.69% of AI-bot traffic versus GPTBot's 9.84%. By May, GPTBot had retaken the lead: 11.48% to ClaudeBot's 9.73%. Then in July, ClaudeBot surged to 16.28% while GPTBot dropped to 9.74% — holding that lead every single day of the month.
Two main forces drive these flips: training cycle intensity and real-time feature rollouts. When a provider enters a heavy model update phase, crawl volume spikes noticeably. When a platform expands its web-access features, its real-time crawler footprint jumps quickly. The July 2026 shift toward ClaudeBot coincides with that platform rolling out expanded web-integrated features.
Worth watching in your own logs: a sustained volume surge from one bot over two or three weeks usually signals a training run or a real-time feature expansion, not random noise.
The ratio that explains the tension
Back to those crawl-to-referral numbers — they put a specific figure on a dynamic that's otherwise hard to quantify.
Training crawls account for roughly 80% of all AI-driven requests hitting your site. Most of that crawling funds model improvements that help AI systems answer questions better. But when a user gets their answer directly in the AI interface, they don't visit your site. Your content contributed to the answer; you saw no traffic.
ClaudeBot was at 286,000:1 in January 2025. By July 2026 it had improved to 2,237:1. GPTBot was at 217:1 in the same period — reflecting the fact that ChatGPT's browse-the-web mode generates meaningful click-through traffic. The platforms doing the heaviest training crawling are not the same ones driving referrals back to sites. That gap is why GPTBot appears in 5.52% of robots.txt DISALLOW rules across the web as of Q1 2026, making it the most-blocked AI crawler by a meaningful margin.
What does any of this mean for your site?
Three practical takeaways from the data.
First: if AI bot traffic is growing in your logs, you're not seeing an anomaly. The sector-wide numbers confirm it's happening everywhere. The variation between sites is in the mix — how much is training versus real-time depends on your content type, your publication frequency, and whether you've been referenced in AI-generated outputs or cited in training data.
Second: the training versus real-time distinction isn't visible in most default analytics setups. Standard log analysis groups all ClaudeBot requests together regardless of intent. You can infer crawl type from behavioural patterns: training crawlers tend to sweep broadly and quickly, hitting many pages in a short window; real-time scrapers are more selective, fetching specific pages in response to queries. Watching request depth, crawl velocity, and which paths get repeatedly hit helps separate them.
Third: the economics are improving, but slowly. The drop from 286,000:1 to 2,237:1 is genuine progress. But 2,237 crawled pages per referred visitor is still a lopsided arrangement if your hosting costs are meaningful. Whether to serve AI bots optimised content, rate-limit heavy training crawlers, or selectively block specific user-agents depends on how that calculation works for your particular site type — and it's a calculation worth actually running rather than ignoring.
The bots are getting smarter about what they crawl and when. The question is whether your approach to managing them is keeping pace.
Sources
- AI Crawler & Bot Traffic Statistics 2026: Key Data — Digital Applied
- GEO Data Report 2026: Crawl-to-Refer Ratio for AI Crawlers — SEOmator
- 2026 State of AI Traffic & Cyberthreat Benchmark Report — HUMAN Security
- Fastly: AI Crawlers Make Up Almost 80% of AI Bot Traffic
- AI Crawler Volume Growth 2022–2026: GPTBot +305% YoY — Presenc AI