The University of Copenhagen announced on September 26 a study finding that AI chatbot answers cover a narrower range of knowledge, and skew toward certain content, compared with the information available through ordinary web search.
The study compared 27 large language models. Newer models generally showed improved diversity of information in their answers, but none reached the level of the web search results used as the benchmark.
That said, this is neither a study of AI accuracy nor a study showing how much knowledge human society has lost because of AI adoption. What it measured is the "breadth of knowledge": how much varied content appears in answers when the same topic is asked about in many different ways.
The research was led by a team including Dustin Wright, now at Aalborg University in Denmark. The paper has been accepted to EMNLP 2026, an international conference on natural language processing. Reading the latest version published on arXiv also shows that the "18.7%" figure used in the university's announcement needs careful interpretation.
About 1.7 million answers broken down into 70 million "claims"
The researchers prepared 200 different question phrasings for each of 155 topics and generated about 1.7 million answers. From these, they extracted about 70 million "claims" for analysis.
The models tested included OpenAI's GPT series as well as the Llama, Gemma, and Qwen series. The topics covered people, historical events, and general concepts related to 12 countries, and the study was designed so the same kinds of questions could be applied across different topics.
A "claim" here refers to an individual piece of explanatory content in an answer.
If the wording differs but the meaning is the same, the claims are grouped together. This makes it possible to distinguish between a model that returns nearly the same explanation every time and one that touches on a variety of content.
For this reason, an answer that is merely longer or more expressive is not considered to have greater knowledge diversity.
The evaluation used the Hill-Shannon diversity index, which is also used in ecology and other fields. It takes into account not only the types of information included but also how frequently each appears.
For example, if four types of content appear in equal proportions, diversity is high. But even if the same four types are present, the value is lower if one of them appears over and over.
This makes it possible to distinguish between a case where a rare item appears only once and a case where a wide range of content appears repeatedly across answers.
The benchmark was a collection of web pages gathered from Google Search results.
On January 1, 2026, the researchers searched each topic name on Google Search for the US and retrieved up to the top 40 pages.
This was not a study comparing Google's AI summary feature with chatbots, nor did it measure how much knowledge people who used search actually acquired.
To keep results from being determined merely by differences in the amount of text collected, the researchers also estimated how well the information distribution had been captured and adjusted sample sizes as needed for comparison.
Even so, the results apply to questions asked in English about topics the researchers chose.
The models compared were also mainly those released from late 2022 through 2025, so the results cannot be read as a ranking covering every AI service available as of 2026.
What does "18.7%" actually compare?
In Section 6.2 of the paper, the diversity averaged across the 155 topics was 3110 for Google Search and 2621 for GPT-5, the highest-scoring of the models tested.
From this, the authors state that search results were at least 18.7% more diverse than GPT-5.
The University of Copenhagen announcement, on the other hand, describes it as GPT-5's information diversity being 18.7% lower than Google Search.
However, the percentage changes depending on which value is used as the baseline.
| Method of comparison | Calculation using the paper's averages | Percentage |
|---|---|---|
| With GPT-5 as the baseline: how much search results exceed it | (3110 − 2621) ÷ 2621 × 100 | About 18.7% |
| With search results as the baseline: how much GPT-5 falls short | (3110 − 2621) ÷ 3110 × 100 | About 15.7% |
With the same figures of 3110 and 2621, using GPT-5 as the baseline gives "search results are about 18.7% higher," while using search results as the baseline gives "GPT-5 is about 15.7% lower."
This table uses the 155-topic averages given in the August 31, 2026 version of the paper, calculated to one decimal place.
The wording closest to the paper itself is therefore: "Google Search was at least 18.7% more diverse than GPT-5."
The phrase "at least" reflects the comparison method.
For some topics, the web pages collected through search did not fully capture the distribution of information that exists. The researchers reduced the search-side sample to equalize conditions only when certain criteria were met.
For this reason, the authors treat the measured advantage of search results as a lower bound.
The figure of about 15.7% can also be obtained by calculating in the opposite direction from the published averages. But this does not mean AI necessarily loses 15.7% of the knowledge that exists on the web.
Lower information diversity is a separate issue from answers being wrong.
Even if answers concentrate on the same content, that content may well be correct. This study alone cannot be used to judge AI accuracy or the rate of misinformation.
Combining web search raises diversity, but still falls short of search results
The study found that using retrieval-augmented generation (RAG), which consults external materials, significantly increased answer diversity compared with a model answering only from its internal knowledge.
In the RAG setup, the top 20 Google Search pages served as sources, and up to 1,000 tokens of question-relevant passages were provided to the model.
Adding external information broadened the range of content appearing in answers.
However, even models using RAG did not reach the diversity of the web search results themselves.
Regarding model size, within the range examined, larger models tended to show lower information diversity.
The authors cite as one possible reason that larger models may more strongly memorize and reproduce information that appears frequently in training data.
This is, however, a hypothesis drawn from the findings; the mechanism was not proven experimentally. Nor does it mean that smaller models are better in overall performance, including accuracy and reasoning ability.
The size of the improvement from RAG also varied by country.
Averaged by country, the more diverse the information in a country's search results themselves, the greater the improvement RAG brought to answer diversity.
The authors also note that using Google Search configured for the US may have influenced regional differences.
In other words, the range of knowledge in an answer depends not only on the AI model itself but also on which sources are searched and passed to the model.
A tendency for English-language information to be reflected more
Regional bias also appeared in an analysis that matched content in English Wikipedia against content in local-language Wikipedias for each country.
In five of the eight countries analyzed, information in English Wikipedia was significantly more likely to be reflected in model answers.
No country showed local-language information reflected significantly more strongly than English.
These results, however, come from questions asked to the models in English. The study does not directly show that the same tendency would appear when using AI chatbots in Japanese.
Being able to explain a region fluently is not necessarily the same as being able to draw widely on the knowledge accumulated in that region.
Even with RAG, if the web pages consulted are skewed toward the same language sphere or similar sources, limits on answer diversity remain.
In evaluating AI answers, what matters is not just whether a system has web search, but also which languages and regions its searched sources come from.
Has "knowledge collapse" already begun?
Another important result of the study is that information diversity tended to improve in newer models.
In every model series except Qwen, diversity metrics tended to rise with each newer generation.
In the university's announcement, Professor Isabelle Augenstein of the University of Copenhagen, a co-author of the paper, explained these improvements and said:
"Knowledge collapse has not happened yet."
What the researchers are wary of is a future risk: as people increasingly get information through AI, society could end up repeating only the limited set of explanations that AI frequently chooses.
Furthermore, if homogeneous AI-generated text proliferates on the web and is reused as training data for the next generation of AI or as a source for RAG, the benefit of broadening answers through external information could weaken.
But the study did not demonstrate that such a cycle has already begun eroding society's knowledge as a whole.
There are also limits in how the research subjects were chosen.
The researchers focused on topics with sufficient Wikipedia coverage, so the study does not extend to extremely rare matters or knowledge shared only within small regions or communities.
The process of breaking AI answers into individual "claims" and grouping those with the same meaning may also introduce errors.
There are regional differences that country-level comparisons cannot capture, and the value of AI answers cannot be judged by a single metric such as diversity.
For users, the study offers a way to think about whether rephrasing a question to the same AI is enough, or whether they need to look into other sources themselves.
For developers, there is room to examine not only accuracy and relevance but also which information is repeatedly selected and which languages' and regions' information is missing from answers.
Combining web search alone is not enough. Only when diverse sources are gathered and their content is actually reflected in the answers can an AI with search capabilities widen the range of knowledge available to its users.
