What is AI crawler visibility?
AI crawler visibility is whether the automated crawlers behind AI answer engines - such as GPTBot, ClaudeBot and PerplexityBot - can find and read a site's pages, which determines whether that content can ever be cited in an AI-generated answer.
Search engines have crawled the web for decades, but a newer generation of bots now crawls it for a different purpose: to train AI models or to fetch a page live when someone asks an AI assistant a question. GPTBot and ClaudeBot are the best known training crawlers; OAI-SearchBot, ChatGPT-User, Claude-User and PerplexityBot fetch pages live, in response to a specific user query, closer in behaviour to a search engine's own crawler. Google-Extended controls whether Google's AI features, separately from classic Googlebot indexing, can use a site's content.
A page cannot be cited in an AI-generated answer if the crawler behind that answer never reads it. Visibility has three distinct layers: a crawler has to find the page at all (discovery, typically via a sitemap or a link), it has to be permitted to fetch it (robots.txt does not block it), and it actually has to request the page content rather than stopping at discovery files. A site can look 'crawled' in server logs while a specific AI crawler has, in practice, only ever touched its sitemap.
This is a young and fast-moving area. There is no universal certification or audited standard yet for 'AI-visible', and any site claiming to guarantee inclusion in an AI answer should be treated with scepticism - no publisher controls whether a model chooses to cite a given page. What a site can measure and improve is the input side: whether the relevant crawlers are reaching its content, and which parts of the site they spend their attention on.
In 99 Data Rooms
How it works here.
99 Data Rooms runs a daily collector against its own site that measures exactly this. It queries Cloudflare's traffic analytics for a defined list of AI and search crawler user agents - including GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Googlebot, bingbot, Applebot, Amazonbot, CCBot and Meta's external agent - and records, per crawler per day, the exact request total plus a breakdown of how many of those requests were discovery paths (robots.txt, sitemaps, llms.txt), how many were static assets (JS, CSS, images), and how many were actual content pages. The results are stored in a permanent table rather than only queried live, because the underlying analytics data expires after eight days. This tells us which AI crawlers visit the site and what they read; it does not and cannot tell us whether any specific AI assistant will cite a specific page, since no site controls that.
Common questions
AI crawler visibility, in short.
Which crawlers count as AI crawlers?
The best known are GPTBot and OAI-SearchBot (OpenAI), ClaudeBot and Claude-User (Anthropic), PerplexityBot (Perplexity), Google-Extended (Google's AI features), CCBot (Common Crawl, used to build many training datasets) and Meta's external agent. Each identifies itself by a distinct user agent string, which is how a site can tell them apart from ordinary search crawlers like Googlebot or bingbot.
Does blocking AI crawlers in robots.txt stop AI systems reading a site?
It stops the crawlers that respect robots.txt, which is most major ones, but it is a voluntary convention rather than an enforced barrier, and it also means the site cannot be read for citation purposes at all. Whether to allow or block AI crawlers is a judgement call - some publishers want the reach, others want to exclude their content from model training.
Does 99 Data Rooms publish live numbers on which AI crawlers visit it?
99 Data Rooms operates a daily collector that measures crawler request counts, split into discovery, content and asset requests, per crawler per day. The measurement exists and is retained as a permanent series; it is not published as a live, continuously updated public dashboard here.
Related terms
What is a virtual data room?
A virtual data room (VDR) is a secure online space for sharing sensitive business documents with outside parties, where every viewer is controlled and every view is tracked.
DefinitionWhat is UK data residency?
Data residency is the physical location where your data is stored and processed; UK data residency means it stays on servers within the United Kingdom rather than being moved abroad.
DefinitionWhat is an audit trail?
An audit trail is a chronological, tamper-evident record of who did what and when to a document or system - a log you can rely on later to prove exactly what happened.
Try it on a real document. Turn a PDF into a tracked, revocable link in a couple of minutes. Three rooms stay free for as long as you want them, no card required.