TLDR: When ChatGPT, Claude, Perplexity, or any other AI assistant conducts web research, it doesn’t search the entire internet, but only a filtered subset of it: an index, a few dozen retrieved pages, and whatever it has been granted access to. This subset has been shrinking measurably for the past two years because websites, publishers, and CDN providers are systematically blocking crawlers. The result still appears comprehensive, is phrased confidently, and is accompanied by sources. That is precisely where the risk lies.
Three layers that even an AI can't tell apart
Anyone who asks an AI a question receives an answer drawn from three very different sources, which blend together in the output in a way that makes them indistinguishable.
The first That's a shift Training Materials. A snapshot of the Web as of a specific date, compiled from corpora such as C4, RefinedWeb, or Common Crawl. The second A shift is the Search Index, which the system accesses, usually Bing, Google, or a custom index. The third A shift is the Live-View: a handful of specific pages that are open and being read at the moment the request is made.
None These three layers comprise “the Internet.” They contain subsets that have emerged according to their own unique rules. And because the language model smooths out the origin of the response text, a statement from a ten-year-old forum post looks exactly the same as one from a primary study accessed today.
That's not a system error. That's just how it's designed.
First bottleneck: Access is regulated
Access to the open web has never been a law of nature, but rather a convention. That convention is currently being abandoned.
The MIT team behind the Data Provenance Initiative monitored 14,000 web domains over the course of a year. Between April 2023 and April 2024, approximately five percent of all tokens in the C4 corpus were completely blocked from AI crawlers via robots.txt. For the most important sources in terms of quality—that is, the most actively maintained domains accounting for the largest share of the training data—the figure exceeded 25 percent. A year earlier, the figure was less than three percent. If terms of service are used as a benchmark instead of robots.txt, 45 percent of C4 is now formally restricted.
The authors state the consequence precisely: If these restrictions are respected or enforced, they distort the diversity, timeliness, and scaling logic of AI systems. They do not make them smaller, but rather more skewed.
Added to this is the infrastructure layer. Cloudflare handles a significant portion of global web traffic and has turned crawler access from a technical issue into a commercial one: one-click blocking, pay-per-crawl, and—starting in 2026—standard blocks for certain bot categories on ad-supported sites. The figures supporting this development are clear. Cloudflare measures the ratio of crawled pages to returned visitors. In July 2025, Anthropic averaged about 38,000 crawls per referral, OpenAI about 1,091, Perplexity 195, and Google 5. Eighty percent of AI crawling was for training purposes, not search.
Cloudflare supplements its access control with „Pay per Crawl“: Website operators can allow or block AI crawlers, or charge a fee for successful content requests. In addition to a standard price per zone, Cloudflare now also supports dynamic pricing, such as pricing based on URL or request characteristics.
From a publisher's perspective, this is no longer a trade-off. From the perspective of your research, it means that more and more doors are closing.
Second Narrowing: What Is Not Visible at All from a Technical Perspective
Even larger than the portion that is intentionally blocked is the portion that a crawler can never reach due to its structure.
Specialized databases are often behind logins, paywalls, and corporate intranets. In addition, there are government systems or curated systems that only return results after form submissions; these include commercial registry queries, standards, and many guidelines that require a fee. There are also archives available as scanned PDFs without text recognition, as well as all applications that load their content only via JavaScript. There are firewalls and rate limits that block bots before they can even view a page.
This so-called "Deep Web" is orders of magnitude larger than the indexable web, and it systematically contains more valuable, more verified, and more specialized knowledge—exactly what a sophisticated research project requires.
In the DACH region, there is an additional language bias. English-language material dominates the corpora. Swiss cantonal law, municipal ordinances, association statistics, German-language scholarly publications, and industry-specific surveys are underrepresented or simply not included. A question about Swiss practice is then answered with a plausible-sounding response derived from U.S. circumstances.
Third Restriction: Retrieval Is Selection, Not Full Access
Even where access is open, a system doesn't read everything. It reads a few results.
A Retrieval-Step formulates a search query, retrieves the top results, extracts text passages, and passes them to the model. A deep research run—which appears impressively thorough—processes anywhere from a few dozen to a few hundred pages, depending on the system. That’s a solid ten minutes’ worth of work. It is not a comprehensive survey.
The selection process follows the ranking logic of the underlying search engine—that is, SEO-optimized pages, aggregated summaries, and syndication platforms. Columbia University’s Tow Center for Digital Journalism tested eight AI search systems with excerpts from real articles, asking for the title, publisher, date, and URL. Over 60 percent of the answers were incorrect. Perplexity performed best with a 37 percent error rate, while Grok 3 had a 94 percent error rate. A common pattern: Instead of the original, the repost on an aggregator platform was cited. Hedging was rare; the incorrect answers were phrased just as definitively as the correct ones.
The major BBC/EBU study from 2025 confirms this picture from a different perspective. Twenty-two public media organizations from 18 countries analyzed over 3,000 AI responses to news-related questions. Forty-five percent contained at least one significant problem, and 31 percent had issues with sources. One detail from the study’s design is more revealing than any percentage: the participating organizations had to temporarily disable their own technical restrictions for the test phase so that the AI assistants could even view their content. Afterward, the restrictions were reactivated.
Fourth Limitation: Contracts Shape Our Perspective
What appears in an AI response depends increasingly on who has a licensing agreement with whom.
OpenAI, Google, Microsoft, and Perplexity have different agreements with various publishers, databases, and platforms. Those with a deal are in the spotlight. Those without one may remain invisible despite having excellent content. This creates an asymmetry that is completely opaque to users.
Practical implication: If you ask the same question to two providers and both give similar answers, that is not confirmation. It may mean that both are drawing on the same subset of available data. And if they give different answers, it is often not due to the model, but to the corpus.
A second opinion from another system is a second perspective, not a verification.
What this means for your research
The answer to this structure is not mistrust, but method.
Check the source instead of taking the quote at face value. If an AI provides a number, click the link and look for the number in the original text. In the studies mentioned, this was precisely where the most common errors occurred.
Ask specifically for the primary source. Don’t ask, “What does the research say about X?” but rather, “Which original study, from what year, with what sample size, and published by which publisher?” Systems that can’t cite a primary source often haven’t seen one.
Check the date carefully. An index can be months old, and a training level can be years old. With everything that changes—such as prices, laws, people, and market figures—the date is more important than the content.
For DACH-related topics, go to the national sources: the Federal Statistical Office, SECO, Fedlex, cantonal portals, and industry associations. This level is often completely missing from global corpora.
And treat every AI response as a hypothesis that requires evidence. The CRAP test remains the simplest framework for this: Currency, Reliability, Authority, Purpose. Four questions that take less than two minutes to answer.
What this means for your website
The same structure works the other way around. If your company’s information isn’t in the index, it doesn’t exist for AI systems—neither for your customers’ searches nor for the response in which you’d like to be mentioned.
This makes the robots.txt file a strategic decision, not just a technical footnote. It’s possible—and usually makes sense—to separate training crawlers from response crawlers: one feeds models, while the other fetches a page because someone just asked about you. A blanket block handles both at the same time.
The opposite is true for internal knowledge. No public model will ever have access to the information contained in contracts, minutes, proposals, and manuals. Anyone who wants to work with this information needs their own retrieval structures based on their own data. Not solely for data protection reasons, but because otherwise there is simply nothing for the system to access.
This shifts the focus of what truly matters. It’s no longer about finding information, but about assessing what’s missing from an answer and why it’s missing. That’s a question no system can answer for you, because it can’t see its own shortcomings.
Key takeaways for you
- AI responses combine three sources: frozen training data, a search index, and a small number of pages retrieved in real time. In the output, these sources are indistinguishable.
- The accessible portion of the web is shrinking measurably. Within a year, over 25 percent of the tokens from C4’s most important source domains were blocked via robots.txt.
- The Deep Web—which includes login-protected areas, paywalls, specialized databases, and government systems—contains more verified information and remains structurally inaccessible to crawlers.
- Retrieval reads results, not holdings. The Tow Center found incorrect source citations in over 60 percent of the tests, often citing secondary sources instead of originals.
- License agreements help determine who is visible. Comparing the responses of two systems is no substitute for verification.
- In the DACH region, there is an additional linguistic and legal bias: national sources are underrepresented.
- Your robots.txt file is a strategic decision about visibility, not just a technical setting.
And that’s why I remain convinced: The future no longer belongs to typing and clicking, but to working together with intelligent systems. But only if we know what these systems focus on—and what they don’t.
Anyone who understands the blind spot in their field of vision works better with the machine than anyone who trusts it blindly. So if you want to talk, and if you want to work with me, feel free to get in touch. www.rogerbasler.ch
Disclaimer
This article was compiled manually based on my own knowledge and supplemented by AI-powered research (Perplexity.Ai and Gemini.Google.com), then simplified using Deepl.com/write. The text is then reviewed and critically evaluated by two individuals of my choosing. The image is from AI-generated imagery (Ideogram/Adobe Firefly). This article is purely educational and does not claim to be exhaustive. The metrics cited regarding crawler behavior change monthly; the values provided refer to the respective measurement period mentioned. Please let me know if you notice any inaccuracies; thank you.
Sources
Deck, A. (March 10, 2025). AI search engines fail to produce accurate citations in over 60% of tests, according to a new Tow Center study. Nieman Journalism Lab. https://www.niemanlab.org/2025/03/ai-search-engines-fail-to-produce-accurate-citations-in-over-60-of-tests-according-to-new-tow-center-study/
European Broadcasting Union & BBC. (2025). News Integrity in AI Assistants: An International PSM Study. EBU. https://www.ebu.ch/research/open/report/news-integrity-in-ai-assistants
European Broadcasting Union & BBC. (2025). News Integrity in AI Assistants Toolkit. EBU. https://www.ebu.ch/files/live/sites/ebu/files/Publications/MIS/open/EBU-MIS-BBC_News_Integrity_in_AI_Assistants_Toolkit_2025.pdf
Longpre, S., Mahari, R., Lee, A., Lund, C., Oderinwale, H., Brannon, W., et al. (2024). Consent in Crisis: The Rapid Decline of the AI Data Commons. arXiv. https://arxiv.org/abs/2407.14933
Tomé, J. (August 29, 2025). The Crawl-to-Click Gap: Cloudflare Data on AI Bots, Training, and Referrals. Cloudflare Blog. https://blog.cloudflare.com/crawlers-click-ai-bots-training/
Cloudflare. (2026). AI Insights. Cloudflare Radar. https://radar.cloudflare.com/ai-insights
Cloudflare. (July 1, 2026). Your site, your rules: new AI traffic options for all customers. Cloudflare Blog. https://blog.cloudflare.com/content-independence-day-ai-options/