ChatGPT Observed Directly Searching Subreddits, Bypassing Bing

Saeed Ashif Ahmed Saeed Ashif Ahmed · · 2 min read

Share this article

ChatGPT was observed directly searching specific subreddits by name and citing their content, challenging previous assumptions about its data sources and Reddit‘s access policies.

This behavior suggests that OpenAI‘s large language model may not exclusively depend on Bing‘s search index for its web results, indicating a potentially independent data acquisition strategy.

During one observed conversation, ChatGPT reportedly retrieved 48 of 71 results and six of eight citations from a single subreddit, r/whatnotapp, over a 3,650-day window, according to findings by researchers including Suganthan Mohanadasan and Ryan Jones.

The direct citation of subreddit content contrasts with another query where ChatGPT fetched numerous Reddit threads but cited none, suggesting that the model’s citation practices are query-dependent, according to the research.

Bing has reportedly ceased ranking Reddit for common commercial queries, further indicating that ChatGPT’s ability to access and cite Reddit content does not solely rely on Bing’s public search index.

Reddit’s robots.txt file, which governs web crawler access, employs IP verification and provides varying responses to different OpenAI user agents, the researchers found. An OAI-SearchBot user agent received a 200 OK status, while GPTBot and ChatGPT-User received a 403 Forbidden status, implying differentiated access levels for OpenAI’s various bots.

These observations challenge the prevailing industry assumption that ChatGPT’s web-based information is directly sourced from Bing’s search results.

The findings suggest that OpenAI may maintain its own independent web index or possess a licensed data feed directly from Reddit, according to Jenny Halasz, another researcher involved in the analysis.

The varied access responses from Reddit’s servers to different OpenAI user agents underscore the complexity of how AI models gather and process real-time information from the internet.

This development could have significant implications for understanding the proprietary data strategies employed by leading AI developers like OpenAI.


Saeed Ashif Ahmed

Written by

Saeed Ashif Ahmed

I’m Saeed, the CTO of Rabbit Rank, with over a decade of experience in Blogging and SEO since 2010. Partner with us to ensure your project is handled with quality and expertise.

Keep reading

Related Articles

Ready to Dominate Search Results?

Let our experts analyze your website and create a custom SEO strategy that drives real results.