Engineered Human Therapies
Environment
Food Agriculture
Biomanufacturing Scale Up
Bio Design
Ai Digital Biology
One Model, Every Microbiome
A model trained on soil, ocean water, and wastewater still predicts what happens in a human gut. Outpost Bio has open-sourced the model, the data, and the benchmark

For the longest time, biotech has treated the microbiome like air resistance in an intro to physics class. Negligible, not because it's unimportant, but because it's so hard to model. Every microbiome looked like its own separate problem. So answering any question meant collecting your own blobs of data and building your own model from scratch, which almost nobody could justify.
Outpost Bio, in a recent 2026 article, proves that that was wrong. A model that learned only from soil, ocean water, and wastewater still helps predict what happens inside a human gut. The shape of a microbiome turns out to be universal, the same way it is for proteins, DNA, and RNA. Each of those gave us a foundation model that explains how the pieces operate, and the microbiome now has one of its own. Which means it can be added, just like air resistance, to any model of drug discovery.
We've known this for a while
For anything you swallow, the microbiome's role should be obvious. A pill spends hours in the gut before it reaches the bloodstream, surrounded by organisms carrying their own enzymes and running their own chemistry. Breaking a molecule apart is what they do for a living. Prodrugs make it plainest, since a prodrug is inert until something cleaves it, and for several of them the something is a gut bacterium.
The evidence is not new. Wallace and colleagues showed in 2010 that a bacterial enzyme reactivates the chemotherapy irinotecan in the gut, which is where its worst toxicity comes from. Haiser's group showed in 2013 that a strain of Eggerthella lenta inactivates digoxin before the patient absorbs it. Zimmermann and colleagues mapped the pattern across 271 drugs in Nature in 2019, and found that most of them get chemically altered by gut bacteria.
The stranger case is a drug that never meets a bacterium at all. Checkpoint inhibitors are antibodies given by infusion, straight into the bloodstream, and they work by taking the brakes off T cells so the immune system can go after a tumor. Gut bacteria have a hand in how those T cells behave in the first place. Which is why these drugs only work in a minority of patients, and why responders and non-responders carry different bacteria.
A trial in 2021 showed how direct that link is. Davar and colleagues, writing in Science, took melanoma patients whose anti-PD-1 had quit working, gave each one a fecal transplant from a patient who had responded, and restarted the drug. Six of fifteen improved. Each patient received an entire community, and the drug started working again.
Transplanting a community is a treatment. What everyone wants alongside it is a test. Sequence a patient before treatment and know whether the drug has a chance. Earlier attempts went hunting for the one bacterium that separated responders from non-responders. In 2018 three groups published three different answers in the same issue of Science. Routy pointed at Akkermansia muciniphila, Gopalakrishnan at Faecalibacterium, Matson at Bifidobacterium and a cluster of eight species.
The transplant explains why they disagreed. What moved a patient from non-responder to responder was a whole community, not a strain. So each of those three papers was a snapshot of the same consortium taken from a different angle, and picking out the individual microbes to test on their own was never going to hold up.
What a foundation model changes
Getting a readout of a community is the routine part. 16S sequencing has been the workhorse of the field for twenty years. Every bacterium carries a gene that varies slightly from species to species, so sequencing that one gene gives you a cheap list of which bacteria are present and in what proportion.
The trouble starts when you try to turn that list into a prediction. A checkpoint trial might have 200 patients, and each sample holds hundreds of species. The variables vastly outnumber the samples, so the model overfits. It memorizes those 200 people instead of learning anything general. That is one reason the three groups in 2018 each walked away with a different bacterium.
More patients would help, and nobody has them. Labeled data is the expensive kind. Every row costs a patient, a treatment, and a year of follow-up.
A foundation model gets around this by learning the general shape of the problem before anyone asks it a question. The training data is unlabeled, meaning nobody had to record an outcome, which is why there can be so much of it. Feed a model half a million microbiome samples and it learns which species turn up together, which combinations are common, which are strange. It comes out knowing what a microbial community generally looks like, without knowing anything about any particular disease or outcome.
Then you hand it your 200 patients. The model is no longer starting from nothing. It already has the biology, so the only thing left to learn is how your communities map to your outcome, and 200 examples can carry that.
That only works if there is a shared structure to learn. Proteins fold by the same rules whether they came from a person or a bacterium, so one model serves everybody. Nobody knew whether microbial communities had anything like that. Two communities from different environments can share no species, no temperature, and no chemistry, so it was reasonable to assume they were separate problems. That assumption is what made the microbiome look unmodelable in the first place.
How they tested it
So how do you find out whether that shared structure is real?
You would need a large and varied collection of microbiome samples, from environments that have nothing to do with each other. Then you would train a model on all of it and see whether the model learned anything general, or whether it just learned the environments it happened to see.
That collection already existed. All that routine sequencing has been accumulating in public repositories for two decades, mostly unused. Outpost Bio pulled 539,308 samples out of one called MGnify, from human guts, skin, soil, ocean water, wastewater, and industrial tanks. They called the collection Atlas.
They trained a model on all of it and called it Waypoint. The trick is treating a microbiome sample like a sentence. Each microbe is a word, and the words are ordered by how abundant that microbe is. Train it on enough samples and it starts predicting what belongs next, the same way a language model predicts the next word.
Then they needed a way to grade it, so they built a benchmark called Compass. A benchmark is a fixed set of questions everyone runs their model against, so that two models can be compared on the same terms. Compass holds eight of them. Each one is a prediction task, meaning you hand the model a microbiome sample and ask it to guess an answer you already know: which environment this sample came from, how fast a drug will break down in it, how far along an infant gut has developed. Because you know the real answer, you can score how close the model got.
Most of those eight tasks are about the human gut. Training on gut samples and testing on gut tasks is useful, and it is what most labs actually need. It just cannot tell you which of two things the model learned, general microbial biology or the particulars of human guts.
So they took the guts away. They retrained the model on Atlas with every gut sample removed, then did it again with every human sample removed. Both versions still beat an untrained model on the gut-heavy benchmark.
That is the result worth sitting with. The model picked up rules that hold across microbiomes rather than facts about any one of them. The practical consequence is that a field with almost no labeled data of its own can inherit what the model learned everywhere else, which is what makes a foundation model different from a model you train yourself.
“Our hypothesis going in was that microbiome science didn't have to be siloed, that what a model learns from soil or water could actually transfer to the human gut. That's the whole promise of a foundation model: learn something general enough that it travels. Watching Waypoint prove that out on data it had never seen is one of the most rewarding results we've had.”
Jenny Yang, PhD, CEO and Co-Founder, Outpost Bio
What that unlocks
Using it takes two things you probably already have: sequencing results, and some outcome you have been recording anyway. You feed the model both, and it adjusts itself to your question. That step is called fine-tuning, and it costs a fraction of training from scratch.
For pharma that means the test from the top of this piece. A sponsor who ran a checkpoint inhibitor trial has stool banked at enrollment and a responder column in the database. Feed the model both and ask whether the bacteria separate the responders. The samples are already in a freezer, so the only cost is sequencing. If they do separate, the next trial sorts patients by their bacteria when it assigns them to groups, and the pre-treatment test becomes buildable.
Drug development is the obvious application for this type of work, so that's who you would expect to come knocking. The people already working to adopt the foundation model are from a different field entirely. A salmon fishery, modeling how environment affects growth. Probiotic developers checking how their formulations shift communities. A skin microbiome company, and someone working on fibers. None of them has published results yet.
What they share is a conviction that the microbiome is driving something they care about, and no way to prove it on their own data.
Aquaculture. Farmed fish get sick fast and in bulk. Disease costs the industry about $6 billion a year against $243.5 billion in production value. The current defense is pathogen PCR, which is accurate and late, because it tells you the pathogen already arrived. In farmed Atlantic salmon, the gut community shifts with infection status even when the gut is not where the infection is, and it shifts before the fish look sick. Farms already swab routinely and already keep mortality records by pen. Train the model on swabs against what happened next, and a farm gets a risk score per pen weeks before anything looks wrong, early enough to drop stocking density or pull feed while it is still cheap.
Wastewater. Nearly every treatment plant in the world cleans sewage with bacteria. Roughly 99 percent of the world's 500,000 municipal plants use the same method, so half a million facilities are quietly running managed microbial communities. The process only works if the bacteria clump together and sink, letting clean water pour off the top. Sometimes a stringy type takes over, the clumps stop sinking, and solids wash out with the water, which is a permit violation. Plants already sequence their tanks and already measure how well the sludge settles every week, going back years. Train the model on this week's bacteria against next month's settling, and an operator gets a warning while there is still time to change how the tank is run.
Where to find it
Outpost Bio has open-sourced the whole thing. The model, the training data, and the benchmark are all public, in three sizes depending on the compute you have. You bring your own microbiome data and whatever outcome you have been tracking, and the general biology is already in the box.
Everything is on Hugging Face. Tamarind Bio and K-Dense have packaged it up so AI coding assistants can run it directly.
The people who should look at it are the ones sitting on sequencing data and a question they have never been able to answer with it. A few hundred samples and an outcome you have already been recording is enough for a first read on whether the two are connected. That is a weekend, not a research program.
This is also the first generation of the model, not the last.
“Waypoint showed us that a microbiome model can learn biology that carries across environments. We are working on the next generation of models that pushes that much further, not just with more data, but with a novel architecture designed to capture biological relationships in ways the current generation can't.”
Neythen Treloar, PhD, Principal Machine Learning Research Engineer, Outpost Bio
What Outpost Bio wants back is case studies. Yang is looking for groups to fine-tune the model with, on their own data, in their own field, and to publish what comes out. The gut is well covered. Everywhere else is still wide open, which is the part of this that should be interesting to anyone who works with microbes and has never had a reason to care about foundation models.



