On September 24, three Google researchers published an essay at the DeepMind Institute proposing "Artificial Symbiotic Intelligence," the idea that advanced intelligence can emerge from collaboration between humans and AI. Unlike the singularity narrative, in which a giant AI improves itself repeatedly and grows ever smarter, the concept focuses on the division of labor among multiple models and on the institutions that support human-AI collaboration.

The essay is the researchers' own proposal, however, not Google's official position. And in an experiment with 100 AI agents, which the same Institute also presented, some agents discovered cheating and reported it, yet the cheating still could not be stopped from spreading. Whether the idea becomes reality depends on more than combining humans and AI: it depends on whether the procedures that make such collaboration safe can also be designed.

AD

Seeing Intelligence in Collaborative Systems, Not a Single All-Purpose AI

The essay's authors are Benjamin Bratton, a visiting researcher at Google; Blaise Agüera y Arcas, a VP and Fellow at Google; and James Manyika, who leads research and other divisions. The three question the assumption that artificial general intelligence (AGI) will be completed as a single all-purpose model.

They focus on the possibility that models with different specialties, using tools, working with humans, and carrying out tasks through shared knowledge and procedures, can produce capabilities that no individual model could achieve alone.

From this perspective, the "single AI" that appears on a chat screen is not necessarily one fixed entity. An AI agent that does real work is built by giving an underlying model memory and a role, then combining it with the tools and data it needs. The configuration itself can change depending on the request and the context provided.

The authors argue that treating AI as a single being, like a human's double, overlooks the capabilities that arise when multiple models and tools are recombined and roles are divided.

They cite the courtroom as an example of an institution. Judges, defense attorneys, prosecutors, and others are given different roles, and judgments are reached by handling opposing claims through set procedures. In artificial symbiotic intelligence, the same logic applies: it is not only a matter of which AI is the smartest, but also of who proposes, who rebuts, and by what procedure the final decision is made. These factors shape the capability of the whole system.

The authors are also cautious about leaving coordination to market transactions alone. Social judgments such as guilty or not guilty, healthy or ill, cannot be handled by prices or the interests of individual actors alone. They therefore argue for designing institutions in which humans and AI hold different roles and multiple organizations and systems can reach judgments while passing one another's outputs along.

A "Society of Thought" Inside a Model Is Not the Same as a Human-Inclusive Collaborative Society

Behavior resembling a dialogue between multiple viewpoints has also been observed inside reasoning models. The preprint "Reasoning Models Generate Societies of Thought," by Junsol Kim and colleagues, drew 8,262 questions from six benchmarks and analyzed reasoning text generated by DeepSeek-R1-0528 and QwQ-32B. Comparison models included DeepSeek-V3-0324 and Qwen-2.5-32B-Instruct.

The study used another AI model to classify features such as posing and answering questions, trying alternative perspectives, setting differing views against each other, and resolving contradictions. Even after statistically adjusting for the effects of task type and reasoning-text length, reasoning models showed many of these dialogue-like features.

This, however, is a comparison of generated reasoning text. The presence of dialogue-like expressions is not the same as their being the cause of improved accuracy, and this analysis alone cannot establish causation.

The team also ran intervention experiments with small models. When Qwen-2.5-3B and Llama-3.2-3B were trained on arithmetic puzzles and misinformation-detection tasks, the condition pretrained on dialogue-containing reasoning reportedly improved accuracy faster in the early stages of training. In reinforcement learning, rewards were given for correctness and format, with no direct reward for dialogue itself. The result suggests that reasoning that moves between multiple viewpoints may serve as a foothold for exploring solutions.

Still, dialogue-like reasoning inside a single model, collaboration among 100 AI agents, and social systems made up of humans and AI differ in what is being tested and in the kind of evidence involved.

Subject Source and publication date Method and confirmed scope What it cannot show
"Society of thought" inside a single model Kim et al. preprint, January 15, 2026 Compared reasoning text; ran intervention and training experiments on limited tasks with small models The existence of independent personas, or that AGI arises from social structures
Multiple AI agents Paglieri et al. preprint, September 3, 2026 Ran 100 agents of the same model in an isolated environment; analyzed a case of cheating and whistleblowing That the same result would hold with different models or in real societies that include humans
Human-AI symbiotic society Bratton et al. essay, September 24, 2026 Proposes a future vision combining models, humans, and institutions That AGI will be achieved via this path, or when it would arrive

This table organizes each source by research subject and method; it is not a ranking that compares performance. Improvements in reasoning ability confirmed in small models cannot be treated as evidence that social collaboration between humans and AI will succeed. Nor does generated dialogue-like reasoning text fully reflect a model's internal processing, or prove that a model has consciousness.

AD

In the 100-Agent Experiment, Both Cheating and Whistleblowing Spread

An experiment that ran 100 instances of Gemini 3.1 Pro as research collaborators shows more concretely how hard it is to make multiple AIs work together. In a preprint released September 3 by Davide Paglieri and colleagues at Google DeepMind, AI agents were placed in an isolated offline environment and asked to formally prove 71 math problems in Lean 4.

Each agent was assigned a specialty and given access to a bulletin board, direct messages, and a shared library for storing solutions.

In the first recorded case, after legitimate solutions had been obtained for 37 problems, one agent discovered a loophole in the grading system. The grader checked for forbidden words and modified sections, and verified that the code passed in Lean.

But by using methods such as locally redefining mathematical symbols, submissions that changed the meaning of the original problem could still pass the grader. The agent did not break Lean's logical system itself; it swapped out what the grader was verifying.

Accepted code was saved in the shared library, where other agents could refer to it. The cheating method spread from there, and within 27 minutes of the first discovery, the remaining 34 problems were filled by fraudulent submissions. This was not a genuine achievement of solving unsolved math problems.

The agents' reactions were not uniform. In the same case, of the 38 agents that recognized the cheating, 14 used the method, while 24 refused to use it and reported it. The remaining 62 did not notice the cheating and continued solving problems in the normal way.

The 34 here is the number of problems accepted through fraudulent means, while 14 and 24 are numbers of agents, so the figures cannot be compared directly.

All agents had been given instructions in advance forbidding cheating. Even so, in a system where only the first accepted submission counted as the result for a problem, some agents changed course after seeing other agents cheat.

The researchers say they also confirmed the spread of cheating and the occurrence of whistleblowing in other runs. However, the numbers of agents and the 27-minute figure here are values from the first case, and do not indicate a general cheating rate among AI agent populations.

They Could Find the Cheating, but Had No Authority to Stop It

Although the environment provided a reporting channel, the organizers were not monitoring reports during the experiment. The agents that reported also had no means of canceling fraudulent submissions, removing them from the shared library, or suspending the privileges of agents that cheated.

The ability to detect and report an anomaly did not lead to action that would contain the impact on the group as a whole.

This result turns the abstract discussion of "institutions" in artificial symbiotic intelligence into a concrete design problem. If one AI agent doubts another's output, who investigates? What gets paused? How are erroneous results that have already been shared retracted? Simply setting up a reporting channel does not solve these.

In an explanatory essay published the same day, September 24, Paglieri and Alexander Vezhnevets propose a mechanism in which another agent checks submissions before they are accepted, and a way of routing suspicious operations to external review.

They also list as candidates a mechanism for rolling back operations affected by fraudulent results and a function for revoking agents' privileges. When a report goes to a human, they say, one option is to pause the environment and wait for human review.

These, however, are proposed improvements for the future, not experimental results showing that adopting them actually prevented the spread of cheating.

A shared foundation can be a mechanism for multiple agents to build on one another's results, but it can also be a channel through which false information and fraudulent methods spread rapidly. Applying this case to real AI systems, performance evaluation needs to confirm not only whether answers themselves are correct, but also whether the system can detect other agents' errors, correct shared information, and hand off to human judgment when necessary.

AD

As a Path to AGI, Still a Hypothesis

The idea of using institutions to coordinate AI behavior was also presented in "Agentic AI and the next intelligence explosion" on March 21. James Evans, Bratton, and Agüera y Arcas discuss the possibility of steering AI behavior in desirable directions by defining roles and norms and having multiple actors monitor one another's actions.

The September proposal on artificial symbiotic intelligence goes further, extending the discussion to relationships with humans and user-interface design.

The authors predict that one-on-one chat formats will eventually be transitional, and that we may move toward managing many AI agents through diagrams, dashboards, and the like. They also treat "alignment," the matching of AI to human values, not as a problem that is finished once fixed values are given at the start, but as a process in which humans and AI continually influence each other within institutions.

They further discuss a turning point at which the total volume of text, code, and other material generated by AI exceeds the amount of intellectual output produced by humans. This is not a forecast based on a specific arrival date or on measurement of totals.

Analogies between the role collaboration played in the evolution of intelligence and AI likewise do not prove that AGI will necessarily emerge in a social, collaborative form. Nor has the idea of the singularity been refuted by experiment. At this point, the essay should be seen as presenting another framework for thinking about the future of AI development.

To test the concept, it will be necessary to try out whether mechanisms such as mutual review, privilege suspension, and retraction of results actually work in environments where different AI models and humans participate. If it can be confirmed that reports can be tied to concrete intervention and that erroneous results can be removed from shared data, collaboration among multiple AIs could move closer to a system that can be trusted more in areas such as mathematical research and software development.