On April 29, 2026, Biohub, the nonprofit science organization backed by Mark Zuckerberg and Priscilla Chan, announced a five-year, $500 million Virtual Biology Initiative. Biohub committed $400 million to generate data and develop measurement technologies, and $100 million to seed a wider international effort. Its goal is a virtual cell: an AI model that predicts how living cells respond to interventions. The Neuron reported that researchers could eventually use such a model to screen ideas on a computer before choosing laboratory experiments.

The launch brought in the Allen Institute, Arc Institute, Broad Institute, Wellcome Sanger Institute, Human Cell Atlas and Human Protein Atlas, with NVIDIA as a technology partner. Biohub said data generated through the initiative would be open and freely available to the scientific community. (Biohub’s launch announcement)
On October 7, Biohub announced an expansion that puts the combined commitment at $1.8 billion. The new money and resources include more than $500 million from the Department of Energy over five years, NIH coordination of datasets developed through more than $500 million in prior federal investment, and $300 million collectively from Google DeepMind, Isomorphic Labs and Meta. (Biohub’s expansion announcement)

The investment changes the scale of the initiative. It does not, on the available record, establish who gets to use the resulting data first.
The $1.8 billion combines different kinds of commitments

The headline total includes funding, data, computation and measurement technology. It is not all new cash. Biohub’s original $500 million is divided between its own data and measurement work and a broader international effort. DOE will invest more than $500 million over five years. NIH will coordinate relevant datasets, repositories and knowledge bases developed through more than $500 million in prior federal investment. The three commercial partners are contributing $300 million together.
A PR Newswire release calls the effort “nearly $2 billion”; Biohub states the total as $1.8 billion. The larger phrase is rounded language, not a separate commitment.
The virtual cell remains a goal. The announcements describe a planned data foundation and expanded measurements, not a completed model or demonstrated accuracy. The ledger available for this account contains no reported model trained on initiative data and no benchmark result. The commitment is substantial; the predictive performance remains untested.
DOE, NIH and Biohub set the data and computing agenda
The DOE Office of Science, NIH and Biohub signed a memorandum of understanding to advance what DOE calls “Super Intelligence,” or SI, in biology. DOE says its National Laboratories bring computational capacity and biological expertise; Biohub contributes SI technologies and biomedical research infrastructure; NIH brings health data, imaging resources and its research enterprise. (DOE’s MOU announcement)
A large dataset becomes useful for model training only when measurements and formats can be interpreted consistently. Biohub says it will work with NIH to standardize datasets for AI training. NVIDIA, already named as a technology partner at launch, will support the work with accelerated computing infrastructure, domain-specific software and technical expertise. (Biohub’s launch announcement)
The language shifts by institution. Biohub frames the effort around AI-ready biological data; DOE’s announcement uses “Super Intelligence in biology” and “SI-ready science.” The releases describe the same collaboration in different terms. That difference matters for how each institution presents the program, but it does not itself indicate a different technical plan.
Biohub says its generated data will be open and free to the research community. The expansion announcement also describes an open resource. Neither source, as represented in the fact record, specifies a release schedule, access queue, licensing details for every dataset, or whether partners receive early access. Those terms matter more to the commercial advantage question than the headline figure alone.
Measurement limits what biological models can learn
A model’s predictions depend on the examples available to it: which cells were measured, which interventions were tested, and under what conditions. Architecture and compute matter too, as do experimental design and data quality. But when a model has not seen a relevant biological response in its training data, a larger system cannot simply recover that missing evidence. The initiative’s stated plan to expand cell-response measurements across more cell types and conditions targets that constraint directly.
This makes the measurement pipeline strategically important. Whoever can generate and interpret useful measurements at scale can support stronger models, although the available facts do not establish that any participant controls the pipeline or its outputs. The three commercial partners are investing $300 million in technologies and multimodal datasets for predictive models. That gives them a significant role in the effort. It does not, by itself, prove that they will receive data before other researchers.
The open-data promise and the timing question can both be real. Public access to the same files would not erase work already done with them, including model development, staffing, compute allocation and publication. Yet without published access terms or a release timetable, claims that the companies will get a head start remain a hypothesis, not a reported fact. The Neuron’s account describes the arrangement as “open data, with a head start for commercial partners.” The initiative’s public materials establish the investment and open-data intention; they do not establish that sequence of access.
My forecast is narrower than the phrase “virtual cell” may suggest. Within 24 months of the October 7, 2026 expansion announcement, I expect peer-reviewed models using initiative data to show useful accuracy on specific perturbation tasks in a small number of cell types. I do not expect a general-purpose system that reliably predicts cellular responses across biology. This prediction can be tested against published papers and benchmarks by October 2028. The meaningful result will be the performance on defined tasks, not the scale of the funding announcement.
Three unresolved issues will show whether the initiative’s openness works in practice. One is governance: NIH is coordinating datasets developed through more than $500 million in prior federal investment, but the public record in the ledger does not establish the terms under which for-profit partners can use them. Another is reproducibility. An independent replication failure would test the models’ claims, though no such failure is reported here. A third is whether five years of planned measurement proves sufficient; any subsequent funding request framed around a data gap would be evidence that the initial effort did not close it. These are watch points, not known outcomes.
Open data still leaves capacity and governance questions
Academic groups without commercial partners may face a resource gap even if they can download the data. Generating measurements, training models and recruiting experienced staff all require capacity beyond file access. The initiative’s scale could make it harder for smaller labs to compete for researchers interested in working at the measurement frontier. That is a risk implied by the scale of the program, not evidence that hiring losses have already occurred.
European research councils face a related coordination challenge if they seek to match the initiative’s measurement throughput and computing resources. The available facts do not show that European programs have lost talent or fallen behind this effort. They do show why “open” alone cannot guarantee equal capacity to use a dataset.
NIH’s contribution also makes governance consequential. The agency will coordinate resources built through prior federal investment, while the commercial partners are investing in the initiative. Whether those firms receive the same data at the same time as other researchers is a question the public announcements do not answer. The arrangement’s value to the companies could come from participation, expertise and timing, but only the first two are visible in the stated commitments. The access sequence remains the key missing term.
The virtual cell depends on evidence before scale
In April, Biohub set out a $500 million, five-year initiative, with $400 million for its own data and measurement work and $100 million for an international effort. By October, the announced commitment had reached $1.8 billion, combining DOE funding, NIH datasets and $300 million from DeepMind, Isomorphic Labs and Meta. The open-data promise remains part of the plan; the expanded partnership adds federal resources and commercial investment.
The virtual cell will ultimately be judged against experiments and benchmarks. The data is promised open, but the announcements cited here leave release timing and partner access terms unspecified. What researchers can predict, and when they can test those predictions, will decide whether the virtual cell becomes a useful tool or remains a funding ambition.