On October 7, 2026, Biohub announced an expansion of its Virtual Biology Initiative, which is building the data needed to train AI models that predict how cells respond. Google DeepMind, Isomorphic Labs and Meta will jointly invest $300 million, and the U.S. Department of Energy (DOE) and the National Institutes of Health (NIH) are also participating. Biohub says the overall collaboration, which combines new funding with existing data, computing resources and other assets, totals $1.8 billion.
Biohub launched the initiative in April. It has since grown from privately led research funding into an effort that also draws on U.S. government laboratory facilities and data infrastructure. The goal is a "virtual cell": a model trained on experimental data that can predict how a cell will respond when something is changed, such as the activity of a gene. Achieving that takes more than observing large numbers of cells. It requires experimental data that lets researchers compare what was changed with what happened as a result.
What the $1.8 billion actually includes
The combined investment from Google DeepMind, Isomorphic Labs and Meta is $300 million, and the share each company contributes has not been disclosed. The resources provided by government agencies differ in nature. The DOE will fund future measurement and computing work, while the NIH will coordinate access to data and other resources built through past research investment.
Biohub's $1.8 billion figure includes the value of existing data and computing resources, not just new funding. Based on the October 7 announcement and the April launch announcement, the contributions and resources of each party break down as follows.
| Party | Announced amount | Timing / period | Main support |
|---|---|---|---|
| Biohub | $500 million | Five-year investment plan announced April 2026 | $400 million for measurement technology and data generation; $100 million for external research support |
| Google DeepMind・Isomorphic Labs・Meta | $300 million combined | October 2026 expansion announcement | Developing foundational technology for predictive models and building datasets that combine multiple measurement methods |
| DOE | More than $500 million | Five years | Investment in experimental measurement, model development, computing and more |
| NIH | Resources built with more than $500 million in past federal investment | Existing resources from past research investment | Coordinating access to datasets, repositories and knowledge infrastructure; working with Biohub to standardize training data |
All figures are in U.S. dollars, but the periods covered and the nature of the resources differ. In particular, the value of the existing NIH resources cannot be treated the same as new cash outlays. The $1.8 billion should be understood as the scale of the overall collaboration as Biohub has presented it.
The DOE's role also goes beyond supplying computing power. Through the "Genesis Mission," it will make national laboratory supercomputers available and combine measurements such as X-ray and neutron scattering and cryo-electron microscopy. It also plans to use autonomous experimental facilities, supporting both cell measurement and AI model training.
Together with the NIH's standardization of existing data, the partnership integrates two efforts: generating new experimental data, and building the infrastructure to use accumulated research results for AI training.
Observing many cells is not enough to predict responses
What Biohub aims to expand is data recording how cells respond to interventions such as genetic manipulation. It will increase the types of cells and experimental conditions covered and also study interactions between cells. The aim is to move from identifying which genes are active in a given cell to predicting what happens when that activity is changed.
For example, suppressing the same gene does not necessarily cause the same change in different types of cells. For AI models to help plan experiments, they need to accurately link cell type and intervention with the responses actually observed. Even if more cells are measured, prediction accuracy under unseen conditions will not necessarily improve if those links remain ambiguous.
Biohub therefore intends to capture not only molecular-level information but also structures and spatial relationships inside cells, and states that change over time.
Of the original $500 million investment plan, the $400 million earmarked for measurement technology and related work includes support for developing cryo-electron tomography, which observes the inside of cells at near-atomic resolution, and microscopes that image millions to billions of cells in living tissue. These are, however, future technology development goals, not measurement capabilities already achieved as of this announcement.
As measurement methods multiply, so does the difficulty of integrating data from different experiments. Biohub says it will establish common data standards and identifiers and build a system that gives access to data through a single portal.
If results are recorded in a unified format showing which cells were measured under what conditions, data from different research facilities become easier to compare, and more of it can be used to train AI models. Investing in new measurement technology and standardizing data for integration are inseparable efforts.
Will predictions hold up for unseen cells?
The challenges facing current cell-response prediction models can be seen in the results of the Virtual Cell Challenge run by the Arc Institute. It is not a test of the effectiveness of Biohub's investment; it is a public competition on how well AI models can predict the changes caused by interventions in cells.
According to Arc's report on the 2025 competition, evaluation used single-cell RNA expression data from about 300,000 cells derived from human embryonic stem cells known as "H1." The data include results from CRISPRi experiments that each suppress one of 300 genes.
Participants trained their models on data from some of the gene-suppression experiments in the same cell line, then predicted how gene expression would change under interventions not included in the training data.
Results varied by evaluation metric. Improvements were seen in metrics that measure how well a model distinguishes responses to different interventions and identifies genes whose expression rises or falls. But on mean absolute error (MAE), which measures the gap between predicted and measured values, Arc reported that nearly all models performed worse than a simple baseline method.
The MAE result alone does not support concluding that predictive models are useless.
Arc explains that measurement noise, missed detection of gene expression and variation between cells all affect the evaluation. Under such conditions, even beating a simple method that uses the average cell state as the prediction is difficult.
The ability to recognize patterns of change caused by an intervention and the ability to accurately predict the magnitude of that change need to be evaluated separately.
The 2026 competition goes a step further, testing whether predictions hold under different cell conditions.
It covers six cell lines derived from different tissues. This time no competition-specific training data is provided; participants receive only gene expression data from control cells without gene suppression and the identifiers of the genes to be suppressed. Post-intervention measurements are withheld and used as the ground truth for scoring predictions.
In other words, participants can see the normal state of the target cells but cannot learn how those cells respond to intervention. They must estimate responses under unseen conditions using knowledge from other cells and experiments.
For researchers who want to study cells that are hard to culture or experiment on, this ability to handle unseen conditions is especially important. However, the 2026 competition is still under way, and the adoption of this evaluation method does not itself show that models can predict well.
A one-year exclusive window for commercial partners
Another point to watch in this initiative is when data generated by commercial partners will be released.
Axios, citing Biohub's chief science officer Alex Rives, reports that commercial partners can use the data they generate exclusively for one year before it is made public. This is described as a mechanism to preserve companies' incentive to provide funding. Axios report
That means that even though Biohub is building a public data infrastructure, not every researcher will be able to use all the data at the same time.
Companies granted early access may be able to start training and evaluating models sooner than researchers waiting for public release. How much difference this head start makes to model performance is not clear at this point.
Researchers considering using the data will need to check which data will be released, when and under what conditions. Even with common standards in place, if the intervention data needed for validation is not accessible, it may not be possible to fully evaluate a model's performance.
Making virtual cells a useful research tool
What Biohub announced is an expanded collaboration to generate data and build research infrastructure. It has not demonstrated the performance of a new general-purpose cell model or any therapeutic effect.
What will shape future model development is how much experimental data can be collected for the cell types and intervention conditions researchers want to study, and how far that data can be integrated with existing data.
As Arc's evaluation approach shows, confirming a model's practical value requires predicting responses to interventions not in the training data and checking the results against measured values.
Whether recognizing response patterns is enough, or whether the magnitude of change must also be predicted accurately, depends on the purpose of the research. Between collecting large amounts of cell data and getting the answers researchers need, concrete performance evaluation like this is essential.
If data from different conditions is released in comparable formats, and outside researchers can check predictions of unseen responses against measured values, virtual cells come closer to being a practical tool for choosing which experiment to run next.
What will determine the success of the $1.8 billion initiative is not how much cell data is collected, but how many experimentally testable predictions can be generated from it.
