On Wednesday October 7, Biohub, the nonprofit research organization backed by Mark Zuckerberg and Priscilla Chan, announced that funding for its Virtual Biology Initiative has reached $1.8 billion. The new backers read like a compute supply chain: the US Department of Energy, the NIH, Google DeepMind, Isomorphic Labs, and Meta. The stated goal is a "virtual cell," an AI model that predicts how living cells respond to interventions.
Here is the part most coverage skipped. Almost none of the $1.8 billion is described as money for training frontier models. According to the joint DOE, NIH, and Biohub announcement, it buys measurements, microscopes, genome sequencing, autonomous laboratories, and standardized datasets. The models come later. This is a bet that biology's binding constraint is not architecture or even compute, but training data that does not exist yet.
Where the $1.8 Billion Actually Goes
The commitments break down along unusually clean lines, at least as far as the partners have disclosed them:
| Contributor | Commitment | What it buys |
|---|---|---|
| US Department of Energy | Over $500M over 5 years | Lab measurement, imaging, modeling, and computation via National Labs facilities |
| Biohub | $500M over 5 years (pledged April) | $400M for technologies like cryo-electron tomography and advanced microscopy, $100M for external research |
| NIH | Coordination of datasets from over $500M of prior federal funding | Repositories, imaging resources, and knowledge bases standardized for AI training |
| Google DeepMind, Meta, Isomorphic Labs | $300M combined | Data generation for the Virtual Biology Initiative |
| NVIDIA | Undisclosed | Accelerated computing infrastructure, software, and expertise |
The initiative also has a long roster of scientific partners: the Allen Institute, Broad Institute, Gladstone Institutes, Human Cell Atlas, Human Protein Atlas, and the Wellcome Sanger Institute, per Pulse 2.0's report. DOE's contribution flows through the Genesis Mission and includes exascale supercomputers, the Joint Genome Institute, the Environmental Molecular Sciences Laboratory, cryo-electron microscopy and tomography, and robotic "autonomous laboratories" that generate and validate data with minimal human handling.
Biohub Head of Science Alex Rives framed the ambition plainly in the announcement: "An accurate predictive model of biology could dramatically accelerate scientific discovery by enabling scientists to perform experiments digitally." The first dataset should be ready in about a year, The Decoder reports.
The Licensing Clause Is the Real Product
The most consequential sentence in the announcement is not about science. It is about exclusivity. According to The Decoder's interview-based reporting, Biohub's Rives confirmed that commercial funders get one year of exclusive access to the data they paid for before it goes public. Government-funded work carries no such restriction.
Run that as an engineer reading a data contract and the design becomes obvious. DeepMind, Meta, and Isomorphic put in $300 million combined. Roughly that buys a twelve-month head start on datasets nobody else on earth can afford to generate. In any domain where the training corpus is the moat, a one-year exclusivity window is the entire game: a model trained on proprietary multimodal cell data in 2027 ships as a product while the public version of that corpus only lands in 2028.
The government money buys the opposite asset. DOE Under Secretary for Science Darío Gil called the partnership "a new standard for open, data-driven science," and DOE's investment explicitly flows through public user facilities. Open data cannot be fenced, so federal dollars are effectively purchasing a commons that no single company can capture. Each side is buying the thing the other's money cannot.
For an open-science advocate the clause is a concession; for a builder it is a roadmap. If you want early access to the most valuable biological dataset ever assembled, the price of admission is now public and roughly $100 million-plus. Everyone else waits twelve months and gets the same bytes for free.
Why This Matters: Biology Is the Small-Data Problem
Every lab in this deal knows the failure mode they are funding against. Language models scaled because the internet had already labeled the data: text that humans wrote, next tokens that were free to sample. Biology has no internet. Measurements come from instruments, one sample at a time, and they are expensive in money, time, and physics. A single cryo-electron tomography run is days of beam time. That is why Biohub's own mission statement says "the sparsity of data in biology" is the barrier, and why Isomorphic Labs President Max Jaderberg, in the Pulse 2.0 announcement quotes, says solving predictive systems biology "requires scaling past the limits of what any single organization can produce today."
As a data scientist, I would read the budget against the scaling-law logic it implicitly accepts. When a domain is data-limited rather than parameter-limited, adding model capacity yields diminishing returns while added measurements yield outsized gains, and the rational spend shifts toward acquisition instruments and standardization, which is precisely the split in the table above: roughly $900M of the disclosed $1.5B in cash sits in measurement and dataset generation rather than compute or model development. NVIDIA's contribution being compute, software, and expertise rather than a headline dollar figure is the tell that even the hardware partner here treats compute as the easy input.
There is a second engineering lesson hiding in the NIH line item. NIH is not primarily writing new checks; it is coordinating datasets that more than $500 million of prior federal funding already produced, with Biohub standardizing them for AI training. Anyone who has built a model on scraped data knows the ratio: maybe 20 percent of the effort is modeling, 80 percent is cleaning, aligning schemas, and deduplicating. This initiative budgets for that reality at national scale. The Genesis Mission data infrastructure projects named in the DOE release, BioDataNexus BRIDGE and the LAMBDA imaging data architecture, are plumbing projects in the best sense: schemas before weights.
The skeptics' case deserves airtime too. Anthropic's own agent-swarm experiment, where Claude flagged a novel enzyme system once and then missed it in all ten reruns, is a reminder that AI biology results can be fragile even when the data exists. A $1.8B corpus does not guarantee a predictive model of a cell any more than a trillion tokens guaranteed reasoning in 2019. The virtual cell is a research program with an invoice, not a product announcement.
One comparison sharpens the bet. When the field ran the ImageNet playbook on biology, protein structure was the domain that finally cracked, and it cracked because a public, curated corpus, the PDB, existed a decade before AlphaFold's architecture caught up with it. Nobody paid a billion dollars to generate those structures; they accumulated through ordinary grants. The explicit premise of this initiative is that cell-state prediction has no PDB, nothing equivalent is accumulating on its own, and instruments plus federal coordination are the only known way to build one on a five-year clock. That premise is falsifiable, and the DOE release even names the mechanism that could make it work: autonomous laboratories under the OPAL project, where robots close the loop between generating a measurement and validating it, cutting the human handling that historically made biological data both slow and irreproducibly heterogeneous.
From a software engineering seat, the autonomous-lab angle deserves more attention than the dollar total. A lab that can propose an experiment, run it, read the instruments, and log results in a fixed schema converts measurement into a pipeline problem, and pipeline problems have known cost curves. If OPAL-style systems hold their throughput targets, the interesting number to watch is not the $1.8B headline but cost per standardized, model-ready observation, because that is what decides whether the corpus keeps compounding after the five-year funding window closes or stalls the day the grants expire.
Outlook: Watch the Dataset, Not the Press Release
Three things to track over the next year. First, the first dataset release, expected in roughly twelve months: its size, modality mix, and license terms will matter more than any model demo until then. Second, whether the one-year exclusivity window holds in practice; open-data coalitions tend to test embargo terms the first time a funder wants to extend them. Third, whether rival labs keep building walls instead of commons. Anthropic has stood up its own biology lab where Claude guides robots through drug experiments, and The Decoder notes the OpenAI Foundation is putting more than $125 million into biological and medical datasets. Two competing data strategies are now live: rent a head start inside Biohub's commons, or own yours outright.
The market read is that Biohub just priced biological training data at roughly a billion dollars of measured instruments per two frontier model generations, and paid for it with a licensing term. If the first dataset is as rich as promised, expect the next round of AI-for-science competition to be fought over measurement capacity and embargo windows, not benchmark scores. The labs that spent October 7 signing this MOU are not racing to build the virtual cell first. They are racing to be the only ones holding the training data when someone finally does.