AI Search Tools Rely on Diverse Data Sources for Functionality

Joyce de Castro Joyce de Castro · · 2 min read

Share this article

Global AI search tools, including Google‘s AI Overviews and Microsoft‘s Copilot, depend on a wide array of data sources, making a comprehensive understanding of these inputs critical for developers and users.

These data sources are categorized into four tiers based on their current and future utility for functions such as grounding, training, and operational actions within AI systems.

Tier 1 sources, considered confirmed and currently active for grounding and actions, include prominent platforms like Google Search, Bing Search, Google Merchant Center, Google Maps, Google Business Profile, Yelp, Wikipedia, Wikimedia, and Reddit, according to industry analysis.

Live publisher pages also fall into the Tier 1 category, indicating their immediate relevance for real-time information and AI-driven actions.

Tier 2 sources encompass data confirmed for current use in training and licensing agreements, while Tier 3 involves confirmed historical data primarily used for pretraining AI models.

Sources with strong evidence or high likelihood of use are classified under Tier 4, reflecting ongoing development and integration.

Key data categories feeding these AI tools span web and search discovery, products and shopping, local information and places, knowledge and reference, and community-driven content.

News and publisher content also form a significant category, with AI developers increasingly seeking licensing agreements.

Google confirmed that its AI Overviews leverage Google Search results for grounding its Gemini model, ensuring relevance and accuracy.

Similarly, Microsoft utilizes Bing search results to power its Copilot AI, integrating current web information into its responses.

In foundational model training, OpenAI‘s GPT-3 and Meta’s LLaMA models have historically used large datasets like Common Crawl for pretraining, providing a vast linguistic and factual base.

OpenAI has also pursued direct content licensing, signing agreements with major publishers such as the Financial Times and Axel Springer to access their content for training purposes, according to company statements.

Other publishers, including News Corp and the Associated Press, have also engaged in licensing discussions or agreements with AI developers, signaling a broader trend in content acquisition.


Joyce de Castro

Written by

Joyce de Castro

Joyce is a core team member at Rabbit Rank and the lead author covering SEO news, algorithm updates, industry trends, and actionable ranking strategies.

Keep reading

Related Articles

Ready to Dominate Search Results?

Let our experts analyze your website and create a custom SEO strategy that drives real results.