
Image credit: Search Engine Journal
Global AI search tools, including Google‘s AI Overviews and Microsoft‘s Copilot, depend on a wide array of data sources, making a comprehensive understanding of these inputs critical for developers and users.
These data sources are categorized into four tiers based on their current and future utility for functions such as grounding, training, and operational actions within AI systems.
Tier 1 sources, considered confirmed and currently active for grounding and actions, include prominent platforms like Google Search, Bing Search, Google Merchant Center, Google Maps, Google Business Profile, Yelp, Wikipedia, Wikimedia, and Reddit, according to industry analysis.
Live publisher pages also fall into the Tier 1 category, indicating their immediate relevance for real-time information and AI-driven actions.
Tier 2 sources encompass data confirmed for current use in training and licensing agreements, while Tier 3 involves confirmed historical data primarily used for pretraining AI models.
Sources with strong evidence or high likelihood of use are classified under Tier 4, reflecting ongoing development and integration.
Key data categories feeding these AI tools span web and search discovery, products and shopping, local information and places, knowledge and reference, and community-driven content.
News and publisher content also form a significant category, with AI developers increasingly seeking licensing agreements.
Google confirmed that its AI Overviews leverage Google Search results for grounding its Gemini model, ensuring relevance and accuracy.
Similarly, Microsoft utilizes Bing search results to power its Copilot AI, integrating current web information into its responses.
In foundational model training, OpenAI‘s GPT-3 and Meta’s LLaMA models have historically used large datasets like Common Crawl for pretraining, providing a vast linguistic and factual base.
OpenAI has also pursued direct content licensing, signing agreements with major publishers such as the Financial Times and Axel Springer to access their content for training purposes, according to company statements.
Other publishers, including News Corp and the Associated Press, have also engaged in licensing discussions or agreements with AI developers, signaling a broader trend in content acquisition.
Source: Search Engine Journal
Written by
Joyce de Castro
Joyce is a core team member at Rabbit Rank and the lead author covering SEO news, algorithm updates, industry trends, and actionable ranking strategies.
Keep reading
Related Articles

ChatGPT Integrates Paid Ads, Organic Recommendations Globally
ChatGPT now blends paid ads with organic recommendations, offering brands new AI visibility strategies and mea...

Amazon Marketing Cloud Offers 5-Year Data for Advertiser Insights
Amazon Marketing Cloud now offers a five-year dataset, providing advertisers with unprecedented insights into...

YouTube adds shopping features to AI search tool, Ask YouTube
YouTube is enhancing its experimental Ask YouTube AI search with new shopping features, including product comp...