
Image credit: Search Engine Journal
Google research indicates that the sequence of subject and object entities in user queries critically influences the recall performance of advanced artificial intelligence models, despite facts being extensively encoded.
The findings highlight a significant bottleneck in accessing stored information within frontier large language models (LLMs), rather than a deficiency in their training data or encoding capabilities, according to a recent paper by Google.
Frontier LLMs, while encoding between 95 percent and 98 percent of facts, often fail to directly recall 26 percent to 34 percent of this information. Recall difficulty notably increases when questions reverse the subject/object entity order compared to how the facts were initially learned during the models’ training phases.
Researchers observed that LLMs could correctly identify facts with reversed subject/object entities in multiple-choice question formats, which suggests the information is indeed present and encoded within the models. This further supports the conclusion that the issue lies in retrieval mechanisms rather than a lack of stored knowledge.
The study found that simply rephrasing questions had an insignificant impact on recall rates. In contrast, the reversal of subject/object order proved to be a critical factor in how effectively models could retrieve information.
Long-tail facts, which represent rare or less frequently encountered information, presented a greater challenge for recall. Google’s research attributed this difficulty to the recall bottleneck rather than an absence of encoding for these specific facts.
Enabling what the researchers termed ‘more thinking’ within LLMs, a process that allows models additional computational steps, improved recall by 40 percent to 65 percent. However, this method is computationally expensive and requires a specific trigger mechanism to activate, the paper stated.
The research also concluded that merely scaling up LLM training is not an effective solution for addressing the identified recall problem. The core issue pertains to information access, not the volume of data processed during training.
While not directly proven by the research, the findings intuitively suggest that optimizing the ordering of subject and object entities in search engine optimization (SEO) content, to align with common user query patterns, could potentially enhance information retrieval, according to the researchers.
Source: Search Engine Journal
Written by
Joyce de Castro
Joyce is a core team member at Rabbit Rank and the lead author covering SEO news, algorithm updates, industry trends, and actionable ranking strategies.
Keep reading
Related Articles

AI Assistants Transform Local Business Discovery, SEO
AI assistants are transforming local search by curating business recommendations. Learn how to optimize your b...

OpenAI’s ChatGPT Bots Access Forbidden Web Pages in Europe
AI bots, including ChatGPT-User, are accessing forbidden web pages despite robots.txt rules. OpenAI justifies...

Cloudflare to block AI bots from sites opting out of data collection
Cloudflare will block AI bots from sites, OpenAI details indexing, Google Analytics adds benchmarks, and legal...