In 2024, a COVID-19 research team measured plasma samples from 3,348 people using mass spectrometry. The raw data alone amounted to 6.4 terabytes. Including the intermediate files needed for analysis, the total swelled to nearly 28 terabytes. In omics research—the field that comprehensively measures biomolecules such as genes, proteins, and metabolites—data on this scale is generated routinely.

To extract meaning from these massive datasets, researchers have turned to generative AI. It predicts protein structures, screens candidate compounds, fills in missing data, and generates synthetic data. Generative AI is a technology that learns patterns and relationships in existing data to create new content; it became widely known through text and image generation, but its applications in biology are now pushing into far deeper territory.

Yet a fundamental problem lurks here. Generative AI produces "hallucinations"—outputs that sound plausible but have no basis in fact. When a chatbot cites a paper that doesn't exist, it might be dismissed as a joke. But when generative AI embedded in a biological research pipeline hallucinates, the consequences are of a qualitatively different order. A nonexistent molecular pattern could enter a paper as a "discovery," and a fictitious disease mechanism could end up shaping treatment decisions.

AD

The Moment a Hallucination Becomes "Evidence"

Thomas Burger, a computational biologist at Grenoble Alpes University (France), has confronted this problem head-on. Burger, a senior scientist at CNRS (the French National Centre for Scientific Research), has specialized in the statistical analysis of mass spectrometry-based proteomics data. In an Opinion paper titled "Keeping generative artificial intelligence reliable in omics biology," published in Cell Press's journal Patterns in January 2026, Burger classified generative AI's biological applications into ten use cases and evaluated the hallucination risk of each.

The core of Burger's analysis is that not all ten applications carry the same risk. The dividing line he draws is clear: is the AI's output a "hypothesis to be verified later through experiments," or is it "data used directly as evidence"? This distinction fundamentally changes the consequences of hallucination.

Hypothesis-Generation Type (Low Risk) Data-Generation Type (High Risk)
Representative examples Drug candidate screening, digital-twin optimization, counterfactual analysis Missing-value imputation, cohort augmentation, data augmentation, bootstrapping
Consequence of hallucination Time and money wasted verifying off-target hypotheses Fictitious biological effects enter papers as "evidence"
Detectability Errors eventually surface because they are tested experimentally Buried as minute distortions within large datasets, hard to detect
Worst-case scenario A promising candidate is mistakenly excluded A nonexistent disease mechanism becomes established as a "discovery"

In hypothesis-generation applications such as screening, AI presents a list of candidate compounds, and researchers select from that list for laboratory testing. Even if the AI selects the wrong candidates, the experimental stage reveals that they "don't work." The loss is limited to time and expense.

Data-generation applications, however, are a different story. When synthetic data is used as a substitute for, or supplement to, experimental data, hallucinations can slip through the verification net. A molecular pattern fabricated by AI enters statistical analysis as an "observed fact," and the conclusions drawn from it get published in papers. At this point, the hallucination is no longer merely a "mistaken prediction"—it has become "evidence supporting a scientific claim."

Signal Distortion Wears the Face of "Discovery"

Burger's remarks to ScienceAlert cut to the heart of the problem: "In most cases, the issue isn't comparing a hallucination side by side with a genuine biological discovery. Rather, it's that during a complex computational process—a generative-AI-assisted workflow—genuine data gets contaminated by hallucination. This is because the workflow converts raw signals acquired through complex biotechnology into biologically plausible descriptions of molecular mechanisms."

What this observation implies is that hallucinations don't necessarily appear as obviously fabricated. Somewhere in the pipeline running from raw mass spectrometry spectra to the final biological interpretation, AI can subtly distort, amplify, or redirect signals. Such a change, even though it could sway the final conclusion, would not be detected as an independent instance of "fabricated data."

"If, during processing, some signals are distorted, amplified, or altered in ways that steer the final biological conclusion in a different direction, it becomes difficult for researchers to notice—unless they have a deep understanding of how the generative AI operated."

AD

The "Hallucinated Structures" Revealed by AlphaFold 3

This problem is not hypothetical. There is already a concrete example.

In 2024, DeepMind's paper on AlphaFold 3, published in Nature (Abramson et al., 2024), reported major advances in predicting the structures of protein-protein interactions and ligand binding. However, the same paper also acknowledged a new problem that arose from the shift to a generative diffusion model. Whereas AlphaFold 2 used a non-generative architecture, AlphaFold 3 adopted a diffusion model. This change caused the model to "invent" plausible-looking three-dimensional structures in intrinsically disordered regions—regions that inherently lack fixed structure—a phenomenon known as "hallucinated structures."

In AlphaFold 2, intrinsically disordered regions were represented with a characteristic ribbon-like appearance, allowing researchers to intuitively judge "there is no structure here." In AlphaFold 3, regions where hallucination occurs can form ordered structures such as alpha helices, making them hard to distinguish by appearance alone. However, the confidence score (pLDDT) tends to fall well below 50 in such regions, offering a warning sign. The development team reported that they substantially suppressed hallucination through "cross-distillation," which incorporates AlphaFold 2's predictions into the training data.

What this case demonstrates is that hallucinations in generative models don't necessarily manifest as clear-cut "errors." In AlphaFold 3's case, there is a supporting indicator in the form of a confidence score. But when generative AI is embedded throughout an entire omics data-processing pipeline, a comparable warning mechanism may not always exist.

Ten Scenarios, Three Categories

Burger organized the ten use cases into three categories.

The first category is hypothesis generation. This includes drug candidate screening, digital-twin optimization of bioreactors, and counterfactual analysis (simulating "what would happen if this gene were knocked out"). Here, hallucinations manifest as "off-target hypotheses" and are eliminated through experimental verification. In Burger's words, "the cost of pursuing a false hypothesis is not much different from following the intuition of a mediocre student, or data from a paper that is later retracted."

The second category is data generation, which includes missing-value imputation, control-group augmentation, data augmentation, bootstrapping, and data preservation. This is the highest-risk domain. Burger particularly emphasizes the danger of generating data for disease groups to supplement cohorts. If hallucination contaminates disease-group data, it can create biased representations of disease and lead to the "discovery" of false biomarkers.

On the other hand, Burger points out that generating data for control groups (healthy subjects) carries comparatively lower risk. If hallucination happens to increase the diversity within the control group, statistical power may decrease slightly, but the risk of false discoveries actually goes down. Moreover, because healthy-subject data has accumulated extensively from past research, it provides a more robust training base for generative AI.

The third category is the improvement of computational biology software. Hallucinations here manifest as degraded software performance, and are said to be detectable through benchmarking.

AD

The Paradox of "Hallucination as Serendipity"

Burger's paper also contains a surprising question: could a hallucination, by sheer accident, lead to a genuine discovery?

When ScienceAlert posed this question, Burger himself admitted he had never considered it before. "I hadn't thought about this before, but I do think it's possible for a hallucination to lead to a genuine discovery."

This idea resonates with the tradition of serendipity in the history of science. In 1928, Alexander Fleming's discovery of penicillin upon returning from vacation arose from what was, in essence, an experimental failure—contamination of a culture dish by mold. Burger is clearly aware of this parallel: "Serendipity has long been recognized. Whether it originates from a generative AI hallucination or from some other wet-lab mishap, ultimately there should be no difference—either from a moral standpoint or from the standpoint of post-hoc verification."

However, this paradox holds only when a hypothesis derived from hallucination is subsequently verified through experiment. A hallucination accepted as a "discovery" without such verification is not serendipity—it is a false discovery.

Hallucinations Are Already Eroding the Scientific Literature

That Burger's proposal is not merely academic speculation is confirmed by the situation in adjacent fields. A large-scale survey posted to arXiv in 2025 (paper number: arXiv:2605.07723) audited 2.5 million papers and 111 million references across arXiv, bioRxiv, SSRN, and PubMed Central. The results showed a sharp rise in nonexistent citations following the spread of LLMs, with a conservative estimate that 146,932 hallucinated citations were mixed into the scientific literature in 2025 alone.

This is text-level hallucination, which differs in nature from data-level hallucination. Nevertheless, both share the same structural problem: the peer review process is not keeping pace with the speed at which hallucinations spread. At NeurIPS 2025, at least 53 out of 5,290 accepted papers (roughly 1%) contained nonexistent citations. Even after review by three to five expert reviewers per paper, these went undetected.

Hallucinations in omics data are likely even harder to detect than fabricated text citations. The existence of a citation can be mechanically verified against a database, but no automated method currently exists to identify a subtle distortion in a mass spectrometry signal as a "hallucination."

Open Questions

Burger's framework offers a starting point for safely integrating generative AI into biological research. But many problems remain unresolved.

First, there is as yet no quantification of the frequency and impact of hallucinations. The paper does not include empirical data showing, for each of the ten use cases, how often hallucinations actually occur and to what degree they affect final conclusions. Burger himself acknowledges that this list is not exhaustive, stating that "undoubtedly, more 'safe' use cases will emerge in the scientific literature going forward."

Second, detection methods have not kept pace. Methods for identifying signal distortion as hallucination, or for distinguishing data processed by generative AI from raw, unprocessed experimental data, have not yet been established. Burger hopes that "researchers will come to apply the same critical eye to generative AI's data that they apply to experimental data," but achieving that will require a transformation in education and culture.

Third, the boundary between hallucination and creativity remains fundamentally unclear in principle. As Burger points out in his paper, the line between an acceptable "creative" output and an unacceptable "hallucinated" one becomes increasingly blurred the further the output strays from average tendencies. Can this boundary be defined in advance, or must it be left to post-hoc verification? As generative AI becomes entrenched as infrastructure for scientific research, the weight of this question will only grow.

A result proposed by AI, no matter how compelling, is not a discovery until it has been independently verified through experiment. This principle is nothing new. But in a world where generative AI is woven deep into the fabric of research pipelines, and the boundary between AI output and experimental data grows increasingly blurred, upholding this principle is becoming a task more difficult than ever before.