We Are Optimizing the Wrong Objective in Biology
The next breakthrough in medicine may not come from sequencing more genomes or folding more proteins. It may come from the enormous corpus of data we already generate every day, and have never once pointed at health.
By Steven Muskal, Ph.D. | The Renaissance Circle | July 25, 2026 | stevenmuskal.com
I have spent nearly four decades on one side or another of the gap between a molecule’s shape and what it does inside a living system. I built structure-based drug design pipelines before anyone called it that. I ran neural networks on protein sequences in the early 1990s, back when reviewers asked, not always kindly, why a chemist was talking about backpropagation. So when I say that AlphaFold is one of the great achievements of modern science, understand that I am not a bystander offering polite applause. I am someone who spent a career failing to do, by hand and by heuristic, what DeepMind’s system now does routinely before lunch.
AlphaFold solved a problem that had genuinely resisted decades of research: given a protein’s amino acid sequence, predict its three-dimensional structure to near-experimental accuracy. That is extraordinary, and nothing here is an attempt to diminish it. But sitting with the result these past several years has left me with a question I cannot set aside, and it is not really a question about proteins at all. Did we point all of this at the right problem?

Nature Was Never Optimizing Structure
Here is the pivot that changed how I think about my own field. Nature does not optimize protein structures. It optimizes reproductive fitness, through an almost unbearably narrow channel: evolution reads and rewrites DNA, and it observes, through selection, whether the organism carrying that DNA survives and reproduces. That is the whole loop. DNA, organism, fitness. Nothing else exists to evolution as a target.
Evolution has no notion of alpha helices, beta sheets, RMSD, or pLDDT. Those are human abstractions, useful ones, that we invented to make an incomprehensible molecular world legible to ourselves. Protein folding, signaling cascades, metabolism, and physiology all persist because they improve fitness, not because evolution was ever aiming at any of them. Everything between the genotype and the outcome is, to the process that shaped it, invisible machinery. We are the ones who opened the box and started naming the gears. And a protein, for the record, does not even hold still: it is an ensemble of shifting conformations exploring an energy landscape, never a single frozen shape. A crystal structure is one useful snapshot under one artificial condition, not the thing itself. I raise that not to relitigate structural biology, which I love, but to make one point: we have poured a generation of talent into predicting a static object that nature never optimized and never stores.
What Language Models Already Taught Us
I did not arrive at this through biology. I arrived at it by watching my other field, artificial intelligence, over the last decade.
For a long time the dominant intuition in natural language processing was decompositional, in exactly the way biology is decompositional today. A machine that truly understood language, the thinking went, would need dedicated modules: one for syntax, one for grammar, one for semantics, one for logic, one for world knowledge. Whole subfields organized themselves around these presumed joints in the problem, each with its own benchmarks and its own architectures.
Then transformer-based language models, trained on nothing more sophisticated than predicting the next token, blew past every one of those hand-engineered systems. Nobody supervised grammar neurons. Nobody hand-labeled a training set for logical inference. One almost embarrassingly simple objective, at sufficient scale, caused syntax, semantics, world knowledge, and something that behaves a great deal like reasoning to emerge on their own, as internal representations, because they were useful for the one thing the model was actually asked to do. Nobody requested them by name.

Biology should sit with that result longer than it has. Perhaps protein structure, pathway membership, and the rest of our beloved intermediate abstractions are not things we should be supervising a model to reproduce at all. Perhaps they should be allowed to emerge, as latent representations, inside a model trained on a broader and more consequential objective, the way grammar emerged inside a language model without anyone asking for it.
The Brain Is a Black Box, and We Use It Anyway
My research director said something to me recently that I have not been able to stop thinking about. Our brains, he pointed out, are black boxes. We do not understand, in any mechanistic and complete way, how they produce a decision, a memory, or a moment of recognition. And yet we rely on them, all day, every day, to do exactly those things. We trust the output without auditing the mechanism.
That is not a failure of science. It is how we operate with nearly every complex system that matters to us. I do not know the control theory that keeps my car on the road, and I drive anyway. We do not need to derive the internals to make use of the behavior. This is worth saying plainly, because medicine has quietly adopted the opposite creed. We have told ourselves that we may only act once we can trace the mechanism, kinase by kinase, receptor by receptor, structure by structure. Understanding mechanism is a noble goal, and I have spent my career chasing it. But it is not a precondition for helping someone. If we treat human physiology the way we already treat the brain, as a black box with inputs and outputs, then the useful question stops being why and becomes which input changes the output for the better.
Our brains are black boxes. We do not know how they work, and we rely on them every single day. Why would we insist on rationalizing every mechanism inside the body before we are allowed to act on it?
The Corpus Biology Refuses to Use
If medicine is a black box of inputs and outputs, then the obvious next question is: what inputs do we actually have? And here is where I think the field has been staring past an ocean of signal while panning for gold in a stream.
We are not, realistically, going to get most of the population genetically sequenced, phenotyped, and enrolled in longitudinal clinical cohorts. That world is expensive, slow, and consented one person at a time. But look at what we already generate, continuously, without anyone having to run a study. Our economy has gone almost entirely electronic. People pay with credit and debit cards. Amazon knows what arrives on the doorstep. Grocery and pharmacy chains know what goes in the cart, and when the pattern changes. Wearables capture heart rate, sleep, movement, and increasingly glucose and blood oxygen, minute by minute. The social platforms, Meta and Amazon chief among them, link products to purchases to behavior to attention at a resolution no hospital system has ever approached.

This is an enormous corpus of human behavior, and it is a physiological record whether we admit it or not. What you buy, how you move, when you sleep, how those patterns drift over months and years: these are outputs of the same organism medicine is trying to help. Diet is purchased. Sedentary decline shows up in step counts and in the shift from groceries to delivery. Alcohol, tobacco, and sugar are line items. Stress and depression change spending, sleep, and social behavior in measurable ways before they ever reach a clinician. We have built, almost by accident, a real-time sensor network over the daily lives of hundreds of millions of people. We just aimed the whole thing at selling them shoes.
We Already Know This Works
The idea that you can infer something private and important from ordinary transactions is not speculative. Entire industries are built on it. Credit scoring takes a sanitized, aggregated trail of financial behavior and produces a number, a FICO score, that predicts default risk well enough to underwrite the economy. Marketers segment neighborhoods down to the zip code and the street, estimating purchasing power, life stage, and household composition from data that never required anyone to fill out a form. These businesses did not first build a mechanistic theory of why a person defaults or buys. They found that the aggregate signal predicts the outcome, and they used it.
The most famous example is almost a parable for what I am describing. More than a decade ago, a large retailer’s statisticians discovered they could identify, from shifts in ordinary purchasing, unscented lotion, certain supplements, a few dozen everyday products, that a customer was very likely pregnant, sometimes before the family had told anyone. The case, reported by Charles Duhigg in The New York Times Magazine, is usually pinned to Target and its statistician Andrew Pole. No ultrasound. No lab test. No medical record. Just the quiet signature of a changing body expressed through a shopping cart. That story is usually told as a privacy cautionary tale, and the caution is warranted. And it is not a relic of a gentler data era. In early 2026 a company reported the same effect on millions of public grocery orders, flagging likely pregnancies weeks ahead from shifts as ordinary as prenatal vitamins and nausea remedies, and offered it, tellingly, as marketing intelligence rather than medicine. But look at the underlying fact: a health state of enormous clinical significance was legible, at scale, in data nobody collected for medicine.
Now generalize it. If a pregnancy is visible in purchasing, what else is? The early drift toward metabolic disease. The behavioral collapse that precedes a depressive episode. The subtle changes in movement and spending that shadow early cognitive decline. The trajectory, not just the snapshot. We could plausibly read health and wellness at the macro level, at the level of a population, a region, a zip code, a street, the same way credit and marketing already read purchasing power. Not to diagnose one named individual against their will, but to see the trends and trajectories of human health in data we are already sitting on.

So I Tested It on Myself
Before I ask anyone to believe this, I tried to break it on the one subject I fully control: myself. For years I have kept a daily record of my own sleep, movement, heart rate, and mood, alongside my card statement and my complete Amazon history. I pointed one at the other, 63 spending series against 41 health measures across 203 days, and went looking for the signal I had just promised you.
I found nothing. Of 630 correlations, exactly zero survived proper correction, even though three quarters of them would have looked significant to the standard test. Every one was noise. That failure does not wound the argument, it sharpens it: one person carries almost no statistical power, and I never claimed a single cart is diagnostic. The signal I believe in lives in the population, not the person, which is also, conveniently, where the privacy risk is lowest.
The Cart Knew Before My Body Did
That null result is about subtle things, and it holds: at the scale of one person, everyday spending told me nothing I could trust. But subtle is not the only setting a body has. In late March 2026 I caught a respiratory infection, my first illness of any kind since 2018. Seven years of nothing, then a fortnight of feeling wretched. An acute event is where ambient data should show if it shows anywhere, and this time it did, across five separate records, and the first one surprised me.
My cart moved before my body did. Between March 28 and April 8 I ordered eleven cold remedies in twelve days, about $167 worth: Vicks VapoRub, a nighttime vaporizing rub, saline nasal mist twice, Throat Coat and Manuka cough drops, cold-and-cough syrup, throat tea, and finally a steam vaporizer. You could pick that basket out of six months of my spending by eye. My respiratory rate, meanwhile, did not cross its normal range until April 1, when it peaked at 4.3 standard deviations above my baseline, four days after that first purchase. The cart led the body by roughly four days. Not because Amazon understands immunology, but because a purchase encodes something no wrist sensor measures: my own judgment that I felt bad enough to act.
My body did confirm it, but narrowly and late. Respiratory rate and resting heart rate both peaked on April 1, and across the acute window my cardiorespiratory signals tracked the buying at about r = 0.6, lagging the first order by two to four days. Everything else stayed quiet. Sleep, sentiment and activity carried nothing; I actually slept more while ill, not less. Heart rate variability, the closest thing I have to a resilience measure, sat near the top of its range going in and fell only after the infection took hold, following the illness rather than warning of it. Across nine years and 3,311 monitored days, no coherent whole-body event ever clears a proper noise floor, and even this infection was too focal to register as one. Nothing physiological crossed a line before I bought the cough syrup. The body confirmed the event. It did not predict it.
It also did not do it alone, and untangling how is its own small lesson. Two devices with nothing in common, an Eight Sleep mattress under me and an Oura ring on my finger, measured my breathing independently and agreed almost exactly: their respiratory rates tracked to within a fraction of a breath per minute, a correlation of 0.92, and both peaked on the same day, April 1. That is real corroboration, not one sensor echoing itself, and it holds for respiratory rate specifically, not for heart rate, where an overnight low and a daytime value are not the same measurement. The ring also recorded resting heart rate every single night, which is exactly why the body line in the figure above is whole, with no gaps to explain. And making sense of even this much required discovering that the label “Apple Health” was hiding a mattress underneath, that the breathing signal was never really my watch, and that three overlapping devices had to be pried apart to see one clean event. A single person’s record is a messy, mislabeled pile. That is not a footnote against the argument, it is the argument: read one timeline and you fight the mess, read millions and it averages out.
A third ordinary record caught the same illness through what I ate. My food log shows intake falling about 21 percent, roughly 2,410 against 3,040 calories a day, but the striking part is composition, not quantity. My daily sugar collapsed from a median near 96 grams to about 10, held there for twelve straight days, and then snapped back above 100 the day I recovered. That snap-back is the single cleanest marker in the whole dataset. Honey appeared on two-thirds of illness days against 12 percent normally. And yet I logged the same number of items and kept the same nine-hour eating window: I did not eat less often, I ate lighter and cleaner. A different record, kept for a different reason, independently drew the same two weeks.
The last record is where honesty costs something. My restaurant habit thinned during the illness, from three or four dinners out a week to about one, and the longest dining gap in my whole record, eight days, sits right on the onset. That looks decisive until you notice that long gaps happen to me about monthly and are usually travel, not illness. This one even begins with a spa weekend, where meals never hit the card as restaurant charges at all. Only after I subtract the trip does a clean signal remain: a second gap from March 30 to April 4, no hotel or spa charges, sitting on the symptom peak, bracketed by a grocery run on March 29. Home and sick, cooking rather than dining. The transaction stream is richer than a sensor precisely because it carries that context, and treacherous for exactly the same reason.
So five ordinary records, most of them never collected for medicine, all caught one illness, and the earliest and cheapest was a shopping cart. I want to be careful about what that is worth. It is n = 1 and a single episode. My purchase dates are not symptom dates, my food log is self-reported, my dining baseline is thin because the card export only begins March 6, and every threshold here was set after I had seen the data, so these are descriptions, not a test. Nothing here generalizes, and that is the point rather than a disclaimer. One person is an anecdote no matter how many instruments you point at him. The signal a health model actually needs does not live in my timeline, or in yours. It lives in the aggregate, in a million of these noisy records laid over each other, which is the one place I cannot look and the platforms already can.
Align the Data We Have, Instead of Stitching the Data We Don’t
This reframing dissolves the problem that has stalled every grand vision of a medical foundation model. The usual plan is to somehow link the disparate medical record systems: reconcile a thousand electronic health record formats, negotiate access, harmonize codes, and fight the privacy and liability battles at every hospital boundary. It is a noble effort and it moves at the speed of institutions, which is to say barely.
The alternative is to start from the corpus that is already digital, already standardized, and already aligned within each platform. Card networks, retailers, wearable makers, and the large consumer platforms each hold clean, longitudinal, machine-readable behavioral data at a scale that dwarfs any clinical dataset. The work is not inventing a new neural architecture; the architectures that learned language are already close to sufficient. The work is alignment: pointing this corpus, in sanitized and aggregated form, at health and wellness outcomes rather than at ad conversion. Just as credit scoring sanitized financial behavior into a number that is useful without exposing every transaction, a health-and-wellness signal could be derived at the population and regional level without ever assembling a surveillance dossier on a named person.
I do not wave away the privacy questions; they are real and they are hard, and the pregnancy story is precisely why they must be answered before this is built, not after. But notice that working at the macro level, on aggregated and consented and sanitized data, sidesteps the very thing that makes the medical-record approach so painful. You do not need to link one person’s genome to one person’s hospital chart to see that a neighborhood’s metabolic health is deteriorating, or that a cohort on a given intervention is recovering faster. The signal that matters most for public health lives in the aggregate, and the aggregate is exactly where the privacy exposure is lowest.
Medicine Is an Intervention Problem
All of this points back to what medicine has always actually been. Medicine is not trying to predict protein structures, or even disease labels. It is trying to answer one question, over and over, for one person or one population at a time: given everything we know, what intervention most improves the future?
Framed that way, the inputs are interventions, small molecules, biologics, nutrition, exercise, sleep, stress reduction, behavior change, and the outputs are trajectories, biomarkers and function and healthspan over time. Everything mechanistic in between, kinase signaling, GPCR pharmacology, protein folding, can remain exactly what grammar was to a language model: a latent representation the model builds only if and when it helps predict the outcome. If a structure-like representation helps a physiology model predict who responds to a drug, it will construct one, in whatever form is most useful, whether or not a crystallographer would recognize it. If it does not help, the model owes us nothing.
The central question stops being “what structure does this protein adopt?” and becomes “given everything we already know about how this person lives, what changes their trajectory for the better?”
The Organism, Not the Structure
I think history will be kind to AlphaFold, but for a reason slightly different from how we celebrate it today. Not as the endpoint of computational biology, the final word on what a protein is, but as the first genuinely convincing proof that a biological representation of real value can emerge from a rich dataset and a powerful model, learned rather than hand-engineered. That is a bigger and more durable contribution than solving structure prediction, because it is a proof of concept for a program far larger than proteins.
The next chapter, I believe, will not be won by folding more proteins or by finally stitching together the world’s medical records. It will be won by whoever first aligns the vast, ignored corpus of everyday human data, the transactions, the wearables, the behavior we already generate, toward health and wellness, and lets a sufficiently powerful model treat the body the way we already treat the brain: as a black box worth acting on, not a mechanism we must first fully explain. We do not need to understand every gear to help the organism. We need to stop insisting that we do.

Steven Muskal, Ph.D. is the CEO of Eidogen-Sertanty, Inc. - a drug discovery informatics company. He has spent four decades working at the intersection of computational biology, AI, and drug discovery. He writes about AI, health, and the intersection of biology and technology at stevenmuskal.com








