
Image credit: Search Engine Journal
ChatGPT was observed directly searching specific subreddits by name and citing their content, challenging previous assumptions about its data sources and Reddit‘s access policies.
This behavior suggests that OpenAI‘s large language model may not exclusively depend on Bing‘s search index for its web results, indicating a potentially independent data acquisition strategy.
During one observed conversation, ChatGPT reportedly retrieved 48 of 71 results and six of eight citations from a single subreddit, r/whatnotapp, over a 3,650-day window, according to findings by researchers including Suganthan Mohanadasan and Ryan Jones.
The direct citation of subreddit content contrasts with another query where ChatGPT fetched numerous Reddit threads but cited none, suggesting that the model’s citation practices are query-dependent, according to the research.
Bing has reportedly ceased ranking Reddit for common commercial queries, further indicating that ChatGPT’s ability to access and cite Reddit content does not solely rely on Bing’s public search index.
Reddit’s robots.txt file, which governs web crawler access, employs IP verification and provides varying responses to different OpenAI user agents, the researchers found. An OAI-SearchBot user agent received a 200 OK status, while GPTBot and ChatGPT-User received a 403 Forbidden status, implying differentiated access levels for OpenAI’s various bots.
These observations challenge the prevailing industry assumption that ChatGPT’s web-based information is directly sourced from Bing’s search results.
The findings suggest that OpenAI may maintain its own independent web index or possess a licensed data feed directly from Reddit, according to Jenny Halasz, another researcher involved in the analysis.
The varied access responses from Reddit’s servers to different OpenAI user agents underscore the complexity of how AI models gather and process real-time information from the internet.
This development could have significant implications for understanding the proprietary data strategies employed by leading AI developers like OpenAI.
Source: Search Engine Journal
Written by
Saeed Ashif Ahmed
I’m Saeed, the CTO of Rabbit Rank, with over a decade of experience in Blogging and SEO since 2010. Partner with us to ensure your project is handled with quality and expertise.
Keep reading
Related Articles

AI leaders urge caution, audits amid rapid technology growth
MIT, Anthropic, and OpenAI issue AI warnings. Businesses must audit AI adoption and prepare for a slower, mess...

Google introduces Search Profile badges amid publisher traffic decline
Google’s new Search Profile badges are seen as a response to declining publisher traffic attributed to the com...

Google updates e-commerce protocol, adds cart transfer, AI insights
Google rolls out UCP updates, enabling cart transfer to merchant sites, expanding checkout testing, launching...