
Image credit: Search Engine Journal
A new automated tool now simplifies the Common Crawl AI Visibility Audit, enabling websites globally to assess their inclusion in AI training datasets.
The online checker, developed by Chris Green and Suganthan Mohanadasan, provides a streamlined method for publishers and developers to understand how their content is accessed by AI crawlers, addressing a critical concern for many digital properties.
Common Crawl, a non-profit organization, collects more than 2 billion webpages monthly, creating free archives that serve as foundational data for most artificial intelligence training models. Approximately 64 percent of large language models (LLMs) developed between 2019 and 2023 utilized Common Crawl’s data, according to the organization.
The CCBot crawler determines a site’s presence within Common Crawl’s extensive archive. Many prominent news organizations, including the BBC, intentionally block AI crawlers like CCBot, with BuzzStream reporting that 75 percent of top publishers in the United States and United Kingdom have implemented such blocks.
Data indicates that approximately 492,000 websites specifically mention CCBot in their robots.txt files, with about 95 percent of these mentions intended to prevent crawling. Cloudflare also began blocking AI crawlers by default on new domains starting July 1, 2025, impacting 3.8 million domains that had not previously configured a robots.txt file.
Previously, Common Crawl had issued an 18-page manual for conducting a manual AI Visibility Audit. The new automated tool digitizes this process, offering a free online checker.
The Common Crawl Visibility Checker features four diagnostic panels: Captures Per Crawl, Robots.txt History, Whose Block It Looks Like, and A Live Probe. These panels provide detailed insights into a website’s interaction with AI crawlers.
The checker also includes fix cards, which offer tailored recommendations for identified issues, and displays sitemap coverage, comparing captured URLs against the site’s sitemap to highlight potential discrepancies.
Organizations like OpenAI and NVIDIA, along with website builders such as Squarespace, are among the entities that rely on comprehensive web data for their AI development and operational needs.
Source: Search Engine Journal
Written by
Joyce de Castro
Joyce is a core team member at Rabbit Rank and the lead author covering SEO news, algorithm updates, industry trends, and actionable ranking strategies.
Keep reading
Related Articles

EU Designates ChatGPT, Reddit, Roblox Under Digital Services Act
The European Commission has designated ChatGPT, Reddit, and Roblox as Very Large Online Platforms or Search En...

Google’s Mueller: Markdown Files Offer No AI SEO Ranking Boost
John Mueller of Google shared his experience regarding Markdown for AI SEO, indicating it does not provide a r...

Google rolls out global AI performance reports, opt-out controls
Google completed the global rollout of generative AI performance reports and opt-out controls in Search Consol...