We Are Optimizing the Wrong Objective in Biology
My shopping cart registered an infection four days before my sensors did. The population evidence now says the same thing at scale: the target was never structure. It was trajectory.
The next breakthrough in medicine may not come from sequencing more genomes or folding more proteins. It may come from the enormous corpus of data we already generate every day, and have never once pointed at health. Increasingly, population-scale evidence agrees.
By Steven Muskal, Ph.D. | The Renaissance Circle | July 26, 2026 | stevenmuskal.com
I have spent nearly four decades on one side or another of the gap between a molecule’s shape and what it does inside a living system. I built structure-based drug design pipelines before anyone called it that. I ran neural networks on protein sequences in the early 1990s, back when reviewers asked, not always kindly, why a chemist was talking about backpropagation. So when I say that AlphaFold is one of the great achievements of modern science, understand that I am not a bystander offering polite applause. I am someone who started my career failing to do, by hand and by heuristic, what DeepMind’s system now does routinely before lunch. In fairness, when I was working in the area, I only had 200 protein structures to work with. The AlphaFold crew (30 years later) had three orders of magnitude - almost 200,000 protein structures and a massive database of sequences and could leverage the concept that varying sequences across species can fold into the same protein structure, but that’s a whole other story. As I’ve said before, content is king - or rather Content Makes Kings.

AlphaFold solved a problem that had genuinely resisted decades of research: given a protein's amino acid sequence, predict its three-dimensional structure to near-experimental accuracy, the result John Jumper and colleagues published in Nature in 2021. That is extraordinary, and nothing here is an attempt to diminish it. But sitting with the result these past several years has left me with a question I cannot set aside, and it is not really a question about proteins at all. Did we point all of this at the right problem?
Nature Was Never Optimizing Structure
Here is the pivot that changed how I think about my own field. Nature does not optimize protein structures. It optimizes reproductive fitness, through an almost unbearably narrow channel: evolution reads and rewrites DNA, and it observes, through selection, whether the organism carrying that DNA survives and reproduces. That is the whole loop. DNA, organism, fitness. Nothing else exists to evolution as a target.
Evolution has no notion of alpha helices, beta sheets, RMSD, or pLDDT. Those are human abstractions, useful ones, that we invented to make an incomprehensible molecular world legible to ourselves. Protein folding, signaling cascades, metabolism, and physiology all persist because they improve fitness, not because evolution was ever aiming at any of them. Everything between the genotype and the outcome is, to the process that shaped it, invisible machinery. We are the ones who opened the box and started naming the gears. And a protein, for the record, does not even hold still: it is an ensemble of shifting conformations exploring an energy landscape, never a single frozen shape. A crystal structure is one useful snapshot under one artificial condition, not the thing itself. I raise that not to relitigate structural biology, which I love, but to make one point: we have poured a generation of talent into predicting a static object that nature never optimized and never stores.

Evolution reads and rewrites DNA, but it has no notion of alpha helices or RMSD, only whether the organism survives to reproduce. Everything in between is invisible machinery.
What Language Models Already Taught Us
I did not arrive at this through biology. I arrived at it by watching my other field, artificial intelligence, over the last decade.
For a long time the dominant intuition in natural language processing was decompositional, in exactly the way biology is decompositional today. A machine that truly understood language, the thinking went, would need dedicated modules: one for syntax, one for grammar, one for semantics, one for logic, one for world knowledge. Whole subfields organized themselves around these presumed joints in the problem, each with its own benchmarks and its own architectures.
Then transformer-based language models, trained on nothing more sophisticated than predicting the next token, blew past every one of those hand-engineered systems. Nobody supervised grammar neurons. Nobody hand-labeled a training set for logical inference. One almost embarrassingly simple objective, at sufficient scale, caused syntax, semantics, world knowledge, and something that behaves a great deal like reasoning to emerge on their own, as internal representations, because they were useful for the one thing the model was actually asked to do. Nobody requested them by name. A 2022 survey of this phenomenon across dozens of model families gave it a name of its own, emergent abilities: capabilities that are simply absent in smaller models and appear, often abruptly, once scale crosses a threshold, unpredictable from the smaller models’ performance curves alone.
No one supervised “grammar neurons.” Syntax, semantics, and reasoning emerged from a single broad objective, next-token prediction, at scale.
Biology should sit with that result longer than it has. Perhaps protein structure, pathway membership, and the rest of our beloved intermediate abstractions are not things we should be supervising a model to reproduce at all. Perhaps they should be allowed to emerge, as latent representations, inside a model trained on a broader and more consequential objective, the way grammar emerged inside a language model without anyone asking for it.
The Brain Is a Black Box, and We Use It Anyway
My UC Berkeley research director said something to me recently that I have not been able to stop thinking about. Our brains, he pointed out, are black boxes. We do not understand, in any mechanistic and complete way, how they produce a decision, a memory, or a moment of recognition. And yet we rely on them, all day, every day, to do exactly those things. We trust the output without auditing the mechanism.
That is not a failure of science. It is how we operate with nearly every complex system that matters to us. I do not know the control theory that keeps my car on the road, and I drive anyway. We do not need to derive the internals to make use of the behavior. This is worth saying plainly, because medicine has quietly adopted the opposite creed. We have told ourselves that we may only act once we can trace the mechanism, kinase by kinase, receptor by receptor, structure by structure. Understanding mechanism is a noble goal, and I have spent my career chasing it. But it is not a precondition for helping someone. If we treat human physiology the way we already treat the brain, as a black box with inputs and outputs, then the useful question stops being why and becomes which input changes the output for the better.
Our brains are black boxes. We do not know how they work, and we rely on them every single day. Why would we insist on rationalizing every mechanism inside the body before we are allowed to act on it?
The Corpus Biology Refuses to Use
If medicine is a black box of inputs and outputs, then the obvious next question is: what inputs do we actually have? And here is where I think the field has been staring past an ocean of signal while panning for gold in a stream.
We are not, realistically, going to get most of the population genetically sequenced, phenotyped, and enrolled in longitudinal clinical cohorts. That world is expensive, slow, and consented one person at a time. But look at what we already generate, continuously, without anyone having to run a study. Our economy has gone almost entirely electronic. People pay with credit and debit cards. Amazon knows what arrives on the doorstep. Grocery and pharmacy chains know what goes in the cart, and when the pattern changes. Wearables capture heart rate, sleep, movement, and increasingly glucose and blood oxygen, minute by minute. The social platforms, Meta and Amazon chief among them, link products to purchases to behavior to attention at a resolution no hospital system has ever approached.

This is an enormous corpus of human behavior, and it is a physiological record whether we admit it or not. What you buy, how you move, when you sleep, how those patterns drift over months and years: these are outputs of the same organism medicine is trying to help. Diet is purchased. Sedentary decline shows up in step counts and in the shift from groceries to delivery. Alcohol, tobacco, and sugar are line items. Stress and depression change spending, sleep, and social behavior in measurable ways before they ever reach a clinician. We have built, almost by accident, a real-time sensor network over the daily lives of hundreds of millions of people. We just aimed the whole thing at selling them shoes.

The Population Evidence Starts to Arrive
When I first wrote about this idea, I could point to the shape of the argument but not to a body of population-scale results behind it. Since then I went looking, specifically for studies with real cohorts, real sensors, and real numbers, not case reports of one or two patients. What I found does not prove the thesis outright, but it is far more than a hunch, and some of it is already at the scale a health foundation model would need.
Start with the closest thing to a direct test of the wearables argument. Jennifer Radin and colleagues at Scripps Research pulled de-identified Fitbit data from 200,000 users across the United States, narrowed to 47,249 people in five states who wore the device consistently for two full years, and compared week-to-week changes in resting heart rate and sleep against the CDC’s official influenza-like-illness (ILI) surveillance numbers. Across more than 13.3 million individual measurements, adding the wearable signal to a standard surveillance model improved the correlation with CDC-reported ILI rates by an average of 0.12 (up to 32.9%), and the final models tracked the official numbers with correlations of 0.84 to 0.97. Nobody enrolled these 47,249 people in a flu study. They just wore a Fitbit, and the exhaust from that ordinary decision turned out to be real-time epidemiological signal.
The same logic holds for a sharper, more acute event: an active infection rather than a seasonal trend. Tejaswini Mishra and colleagues at Stanford followed nearly 5,300 people wearing consumer smartwatches and identified 32 who tested positive for COVID-19 during the study window. Twenty-six of those 32 (81%) showed a detectable alteration in heart rate, step count, or sleep. Of the cases with both a physiological signal and known symptom-onset timing, the alteration appeared at or before symptoms in the large majority, and four cases were flagged at least nine days before the person felt sick enough to notice. Using a simple two-tier alerting rule built only from resting-heart-rate deviation, the authors estimate that 63% of cases could have been caught before symptom onset, in real time, from a device already on the wrist. A companion effort from the Scripps DETECT study, run by Giorgio Quer, Eric Topol, and colleagues, enrolled over 30,000 participants and found that combining self-reported symptoms with sensor data discriminated COVID-positive from COVID-negative symptomatic people with an AUC of 0.80, significantly better than symptoms alone.

Now widen the lens from wearables to the purchases themselves, which is the part of my own argument I could least support with someone else’s data until recently. Riccardo Di Clemente and colleagues, working with anonymized credit card transaction sequences from an entire metropolitan population, applied a text-compression technique borrowed from information theory to the order in which people buy things, not just what they buy. The method is worth pausing on, because it is the same move that made language models work. They treat each recurring pair of consecutive purchase types as a “word,” compress each cardholder’s history into the vocabulary of words they actually use, and then measure how similar any two people’s vocabularies are. No categories are imposed in advance. The structure has to fall out of the sequence itself.

What fell out were groups distinctive enough that the authors could name them: Commuter, Household, Young, Hi-Tech, and Dinner-out, alongside a large “Average” group of about 39% with no strong signature. Then came the check that matters. When the authors compared those purchase-derived groups against independent data, age, total expenditure, gender, and the diversity of each person’s mobility and social network, drawn from separate mobile-phone records, the groups lined up with all of it. A shopping history, read as a sequence rather than a snapshot, is already legible as a lifestyle. Health is one part of a lifestyle. It is not a large leap from there to a purchase-sequence signature of the kind my own cart displayed the week I got sick, just averaged over a metropolitan population instead of one household.

These three results, wearables tracking flu at the state level, smartwatches flagging infection before symptoms, and purchase sequences legible as lifestyle, are not the same study asked three times. They are three independent groups, three different data types, and three different outcomes, converging on the same underlying claim: ordinary, ambient, already-collected behavioral data carries a real and extractable physiological signal at the population scale. That is precisely the corpus this article is arguing we have been sitting on.
We Already Know This Works
The idea that you can infer something private and important from ordinary transactions is not speculative. Entire industries are built on it. Credit scoring takes a sanitized, aggregated trail of financial behavior and produces a number, a FICO score, that predicts default risk well enough to underwrite the economy. Marketers segment neighborhoods down to the zip code and the street, estimating purchasing power, life stage, and household composition from data that never required anyone to fill out a form. These businesses did not first build a mechanistic theory of why a person defaults or buys. They found that the aggregate signal predicts the outcome, and they used it. Di Clemente’s lifestyle groups are the same move, applied to health-adjacent behavior instead of credit risk.
The most famous example is almost a parable for what I am describing. More than a decade ago, a large retailer’s statisticians discovered they could identify, from shifts in ordinary purchasing, unscented lotion, certain supplements, a few dozen everyday products, that a customer was very likely pregnant, sometimes before the family had told anyone, a story Charles Duhigg reported in the New York Times Magazine in 2012. No ultrasound. No lab test. No medical record. Just the quiet signature of a changing body expressed through a shopping cart. That story is usually told as a privacy cautionary tale, and the caution is warranted. But look at the underlying fact: a health state of enormous clinical significance was legible, at scale, in data nobody collected for medicine.
Now generalize it, and we no longer have to generalize purely on faith. If a pregnancy is visible in purchasing, and a purchase sequence is visible as a lifestyle, and a lifestyle correlates with age and social network, then the early drift toward metabolic disease, the behavioral collapse that precedes a depressive episode, and the subtle changes in movement and spending that shadow early cognitive decline are exactly the kind of signal this new literature says should be there. We could plausibly read health and wellness at the macro level, at the level of a population, a region, a zip code, a street, the same way credit and marketing already read purchasing power, and the same way Radin's team already reads flu season from a wristband. Not to diagnose one named individual against their will, but to see the trends and trajectories of human health in data we are already sitting on.
So I Tested It on Myself
Before I ask anyone to believe this, I tried to break it on the one subject I fully control: myself. For years I have kept a daily record of my own sleep, movement, heart rate, and mood, alongside my card statement and my complete Amazon history. I pointed one at the other, 63 spending series against 41 health measures across 203 days, and went looking for the signal I had just promised you.
I found no detectable relationship. Of 630 correlations, exactly zero survived proper correction, even though three quarters of them would have looked significant to the standard test. Every one was noise. That failure does not wound the argument, it sharpens it: one person carries almost no statistical power, and I never claimed a single cart is diagnostic. The signal I believe in, and the signal Radin, Mishra, Quer, and Di Clemente all found, lives in the population, not the person, which is also, conveniently, where the privacy risk is lowest.
My Cart Beat My Sensors, and I Beat My Cart
That null result concerns subtle, everyday things, and it holds: at the scale of one person, ordinary spending carried no detectable relationship to my health. But subtle is not the only setting a body has. On March 24, 2026, I woke with the first symptoms of a respiratory infection, my first illness of any kind since 2018. Seven years of nothing, then about two and a half weeks of feeling wretched. An acute event is where ambient data should show itself if it shows anywhere, and this time it did, across four separate records. What surprised me was the order in which they arrived.
Four days after I felt it, on March 28, I ordered four cold remedies. Seven more followed over the next eleven days, eleven orders in twelve days for about $167: Vicks VapoRub, a nighttime vaporizing rub, saline nasal mist twice, Throat Coat and Manuka cough drops, Boiron cold-and-cough syrup, throat tea twice, and finally a steam vaporizer. You could pick that basket out of six months of my spending by eye. My wearable signals, meanwhile, stayed inside their normal range until March 29 and did not peak until April 1, when my respiratory rate hit 4.3 standard deviations above my personal baseline. Symptoms, then cart, then sensor, each step about four days after the last.
I want to be precise about what that ordering means, because the tempting version of this story is the wrong one. The cart did not know first. I did. My own felt symptoms led my purchases by four days, and my purchases led my sensors by another four. What the cart did was record a piece of knowledge no wrist sensor can measure: my own judgment that I felt bad enough to spend money on it. That is not Amazon understanding immunology. It is an ordinary behavioral record capturing subjective state, timestamped and machine-readable, four days before any physiology crossed a threshold. For a model trying to read health from ambient data, that distinction is the entire opportunity. The earliest and cheapest signal in my record was not a measurement of my body. It was a decision I made about my body.

My body did respond, clearly and measurably. This was not a subtle event lost in the noise. Respiratory rate reached 4.3 standard deviations above baseline, heart rate variability fell to about 1.7 standard deviations below it at onset before recovering, and resting heart rate rose alongside both. Across the acute window each of those signals tracked the remedy purchases at about r = 0.6, peaking two to four days after the first order. Everything else stayed quiet. Sleep, sentiment, and activity carried nothing for this event, and I actually slept more while ill, not less.
But nothing predicted it. Heart rate variability, the closest thing I have to a resilience measure, sat near the top of its range going into the illness and dropped only after the infection had taken hold. There was no dip beforehand that could have forecast anything, and nothing physiological crossed a line before I bought the cough syrup. My body confirmed the event. It never anticipated it. That deserves saying plainly, because anticipation is exactly what consumer wearables are usually sold on.
It also did not confirm the event alone. Two devices with nothing in common measured my breathing independently: an Eight Sleep mattress underneath me and an Oura ring on my finger. They share no hardware and no data, and they agreed almost exactly, tracking to within a fraction of a breath per minute across the illness, a correlation of 0.90, both peaking on the same day, April 1. That is genuine cross-device corroboration rather than one sensor echoing itself. The agreement holds for respiratory rate specifically and not for heart rate, where the ring’s overnight low and Apple’s daytime value are simply not the same measurement and should never be treated as one.

Now the part that is least flattering and most useful. Getting to that clean two-line chart meant discovering that the label “Apple Health” had a mattress hiding underneath it. The respiratory signal in that export is overwhelmingly the Eight Sleep, not the Apple Watch, which barely measured breathing here at all; the watch’s real contribution was heart rate variability and daytime resting heart rate. The resting-heart-rate line in the cart figure is the ring’s overnight measurement, chosen because it recorded every single night, including April 2 and April 6, the two nights Apple’s separate daytime metric was never computed at all. Even the mattress has holes: it recorded nothing across a spa trip in mid-March and an overnight in late March, both comfortably before symptoms began, which is why its line is dotted there and why the ring had to cover for it. Three overlapping devices had to be pried apart, and each one’s blind spots filled by another, before one clean event was visible.
That messiness is not a footnote against the argument. It is the argument. A single person’s record is a pile of overlapping, gappy, mislabeled instruments, and no one timeline in it is complete. What rescued this event was not a better sensor. It was a second, independent one that happened to be recording when the first was not. That is the aggregate thesis at the smallest possible scale, two devices instead of one, and it is exactly what Radin’s team was doing on a different order of magnitude when it turned 47,249 Fitbits into a signal cleaner than any single wristband could ever offer. Read one timeline and you spend your time fighting the mess. Read millions and the mess averages out.
A fourth ordinary record caught the same illness through what I ate. My daily calories fell about 25% below the two weeks preceding symptoms, a median of 2,408 against 3,207, and recovered within a week. The sharper marker was not quantity but a single ingredient. Honey, the oldest sore-throat remedy there is, appeared more often and in larger amounts as the illness ran, peaked at four servings on April 3, and stopped entirely once I was better. My rhythm never changed at all: I logged the same dozen or so items a day throughout, in the same eating window. Only the contents shifted.
An earlier version of this section claimed something stronger, and it was wrong in a way worth showing rather than quietly deleting. I had reported that my daily sugar collapsed from a median near 96 grams to about 10, held there for twelve days, and snapped back the day I recovered. It did not. My food log did not reliably estimate sugar in grams until early April: only 3% of pre-illness food items carried a sugar value, against 99% afterward. The collapse was missing data wearing the costume of a finding, and it was the cleanest-looking result I had. I caught it only by going back and asking what the denominator was. Honey servings, read from the meal descriptions themselves rather than from an inferred nutrition estimate, survived that check. Sugar in grams did not, so it is not charted here and no longer appears in this essay.

So four ordinary records, two of them never collected for medicine, all registered one illness, and the earliest of them was a shopping cart. I want to be careful about what that is worth. It is n = 1 and a single episode. My purchase dates are not symptom dates, and my symptom dates are not infection dates. My food log is self-reported. Every threshold here was set after I had already seen the data, so these are descriptions rather than tests. And not one of these records spoke until I had stripped a confound out of it by hand, whether that was a travel gap in a sensor or a missing denominator in a food log. Nothing here generalizes, and that is the point rather than the disclaimer. One person is an anecdote no matter how many instruments you point at him. The signal a health model actually needs does not live in my timeline, or in yours. It lives in the aggregate, in a million of these noisy records laid over each other, which is the one place I cannot look and the platforms, and now a handful of academic teams, already can.
A Necessary Caution
I would be doing this argument a disservice if I only told the flattering half of the population story. In 2009, Jeremy Ginsberg and colleagues at Google showed that the frequency of certain search queries tracked CDC influenza data closely enough to estimate flu activity roughly a day faster than official reporting, a genuinely elegant early proof that ambient digital behavior encodes health signal. Google Flu Trends became the flagship example of exactly the kind of system I am describing here. It also broke. By the 2012-13 season, the system was overestimating peak flu prevalence by roughly a factor of two, and in 2014 David Lazer, Ryan Kennedy, Gary King, and Alessandro Vespignani published a widely cited postmortem, “The Parable of Google Flu”, documenting how algorithm changes, media-driven search behavior, and the absence of any updating against ground truth let the model drift quietly away from reality for years before anyone noticed.
I think that failure and my own null result are teaching the same lesson from opposite directions, and my retracted sugar collapse is the same lesson a third time, at household scale. My 630 correlations produced nothing because one household has no statistical power. Google Flu Trends produced a confident, wrong answer because it had enormous scale but no discipline, no correction for its own drift, no ground truth in the loop. And my sugar result looked like the cleanest finding in my dataset right up until I asked how many of the underlying items actually carried a measurement. The lesson is not that ambient data is unreliable. It is that ambient data is exactly as reliable as the statistical rigor applied to it, and no more. Radin's team validated continuously against CDC numbers. Mishra's and Quer's teams validated against confirmed test results. Any health model built on transactions and wearables has to be held to that same standard, recalibrated against ground truth constantly, or it will eventually become a more sophisticated version of the same mistake, at a scale that matters far more than mine.
Align the Data We Have, Instead of Stitching the Data We Don’t
This reframing dissolves the problem that has stalled every grand vision of a medical foundation model. The usual plan is to somehow link the disparate medical record systems: reconcile a thousand electronic health record formats, negotiate access, harmonize codes, and fight the privacy and liability battles at every hospital boundary. It is a noble effort and it moves at the speed of institutions, which is to say barely.

The alternative is to start from the corpus that is already digital, already standardized, and already aligned within each platform. Card networks, retailers, wearable makers, and the large consumer platforms each hold clean, longitudinal, machine-readable behavioral data at a scale that dwarfs any clinical dataset, and, as of the last few years, so do a growing number of academic research groups. The work is not inventing a new neural architecture; recent proposals for generalist medical foundation models argue that the multimodal architectures behind today’s large language and vision models are already close to sufficient for this purpose. The work is alignment: pointing this corpus, in sanitized and aggregated form, at health and wellness outcomes rather than at ad conversion. Just as credit scoring sanitized financial behavior into a number that is useful without exposing every transaction, a health-and-wellness signal could be derived at the population and regional level without ever assembling a surveillance dossier on a named person.
I do not wave away the privacy questions; they are real and they are hard, and the pregnancy story is precisely why they must be answered before this is built, not after. But notice that working at the macro level, on aggregated and consented and sanitized data, sidesteps the very thing that makes the medical-record approach so painful. You do not need to link one person’s genome to one person’s hospital chart to see that a neighborhood’s metabolic health is deteriorating, or that a cohort on a given intervention is recovering faster. The signal that matters most for public health lives in the aggregate, and the aggregate is exactly where the privacy exposure is lowest, and, per Radin and Di Clemente, exactly where the signal has already been shown to hold up.
Medicine Is an Intervention Problem
All of this points back to what medicine has always actually been. Medicine is not trying to predict protein structures, or even disease labels. It is trying to answer one question, over and over, for one person or one population at a time: given everything we know, what intervention most improves the future?
Framed that way, the inputs are interventions, small molecules, biologics, nutrition, exercise, sleep, stress reduction, behavior change, and the outputs are trajectories, biomarkers and function and healthspan over time. Everything mechanistic in between, kinase signaling, GPCR pharmacology, protein folding, can remain exactly what grammar was to a language model: a latent representation the model builds only if and when it helps predict the outcome. If a structure-like representation helps a physiology model predict who responds to a drug, it will construct one, in whatever form is most useful, whether or not a crystallographer would recognize it. If it does not help, the model owes us nothing.
The central question stops being “what structure does this protein adopt?” and becomes “given everything we already know about how this person lives, what changes their trajectory for the better?”
The Organism, Not the Structure
I think history will be kind to AlphaFold, but for a reason slightly different from how we celebrate it today. Not as the endpoint of computational biology, the final word on what a protein is, but as the first genuinely convincing proof that a biological representation of real value can emerge from a rich dataset and a powerful model, learned rather than hand-engineered. That is a bigger and more durable contribution than solving structure prediction, because it is a proof of concept for a program far larger than proteins.
What has changed since I first made this argument is that the proof of concept is no longer confined to proteins and language models. Wearables tracking flu at the state level, smartwatches catching infection before it is felt, purchase sequences legible as lifestyle, and one household’s shopping cart outrunning its own wrist sensors by four days, all of it points the same direction. The next chapter, I believe, will not be won by folding more proteins or by finally stitching together the world’s medical records. It will be won by whoever first aligns the vast, ignored corpus of everyday human data, the transactions, the wearables, the behavior we already generate, toward health and wellness, held to the same discipline that made Radin’s and Mishra’s results credible and that Google Flu Trends lacked, and lets a sufficiently powerful model treat the body the way we already treat the brain: as a black box worth acting on, not a mechanism we must first fully explain. We do not need to understand every gear to help the organism. We need to stop insisting that we do.
Steven Muskal, Ph.D. is the CEO of Eidogen-Sertanty, Inc. - a drug discovery informatics company. He has spent four decades working at the intersection of computational biology, AI, and drug discovery. He writes about AI, health, and the intersection of biology and technology at stevenmuskal.com







I was thinking of your article as I read this in the NYT
https://www.nytimes.com/interactive/2026/07/29/magazine/inflammation-chronic-immune-system-health.html?campaign_id=190&emc=edit_ufn_20260801&instance_id=179685&nl=from-the-times®i_id=76257506&segment_id=224041&user_id=0a49ffa1798023044714530743be5e92
Potential for new readouts for chronic inflammation, emerging as a 'master marker' for long-term disease. Could be very interesting as part of your correlation analysis.
Also note the reference therein to Eric Topol's book 'Super Agers', looks to be in your wheelhouse.
Steve, this was a well-written and thought-provoking article. It will take me some time to fully digest. How does self-selection bias affect this? Some people (my wife) are devotees of wearables and data tracking; others (me) have no wearables and could care less. If this results in systematic issue at a population level, it seems like it could be an issue. A similar phenomenon may hold across other parameters, like purchasing, depending on people's behavior (takes me a long time to admit that I need therapy for a cold). At a minimum, seems like it would increase the noise, and possibly lead to false correlates. Anyway, thanks - write that book!