The Average Patient Doesn’t Exist
What you will learn in this article
I have spent much of my career with clinical trial data: reading it, analyzing it, and explaining it to doctors. And for years one question kept bothering me. If a trial clearly shows that Drug A works better than Drug B, why do so many experienced doctors still say, “It depends on the patient”?
To me, it looked like a convenient excuse. Data is data. If the numbers say A beats B, then prescribe A. What’s the point of a 10,000-patient study if every doctor ends up doing their own thing anyway?
Then a 2026 paper on inflammatory bowel disease changed how I think about this. In this article, I’ll walk you through that paper in plain language. I’ll also show you why anyone who has ever compared two mutual funds already understands the core problem, even without knowing it. By the end, you’ll see why clinical trials are excellent for making decisions about populations, why they are surprisingly weak at making decisions about one person, and what questions you should ask the next time a doctor prescribes you something.
I’ll start with an imaginary exercise from your world. Then I’ll introduce the paper and what it found. After that, I’ll explain why the findings matter well beyond gut disease, and I’ll end with a real story about a heart drug that sums up everything.

Two Funds and a Very Tempting Subtraction
Let me ask you to use your imagination a little, because the whole idea in this article becomes simple once you see it through a finance lens.
Imagine your relationship manager hands you factsheets for two funds. Fund A beat its benchmark by 26%. Fund B beat its benchmark by 9%. Any first-year analyst would subtract the two and conclude that Fund A is about 17 points better. Decision done!
Now read the fine print (which, of course, nobody does).
Fund A’s track record covers 2008 to 2012. Fund B’s covers 2006 to 2010. They were measured against different benchmarks. Fund B’s portfolio had far more defensive holdings. And here’s the killer: Fund A’s 26% only counts the positions that were still doing well after the first six weeks. Everything that underperformed early was quietly dropped before the performance clock started ticking.
Would you still trust that 17-point gap? Of course not. You’d call it what it is: apples versus oranges, with a heavy dose of survivorship bias.
Keep this picture in mind. Because, believe it or not, this is very close to how doctors and analysts have often compared medicines.

A Quick Primer: How Drugs Get Tested
Before we get to the paper, let me set some context. No need to remember the technical terms here.
When a new drug is developed, it goes through a large trial called a phase 3 trial. Patients are randomly split into two groups. One group gets the drug, and the other gets a placebo (a dummy pill or injection with no active ingredient). At the end, researchers compare how many patients improved in each group. The difference is called the effect size. If 40% improved on the drug and 15% on placebo, the effect size is 25 percentage points.
This is the gold standard for proving that a drug works. Regulators across the world rely on it to approve medicines.
But notice what a placebo-controlled trial does not tell you. It doesn’t tell you whether Drug A is better than Drug B. For that, you need a head-to-head trial, where patients are randomly assigned to Drug A or Drug B and compared directly.
Point is, head-to-head trials are rare. They are hugely expensive, take years, and are hard to recruit for. And with so many drugs available today, there will never be enough of them to cover every possible comparison.
So what do people do? Exactly what you did with the two funds. They take the effect size of Drug A versus placebo, subtract the effect size of Drug B versus placebo, and call the difference a ranking.
The Paper That Tested the Shortcut
(The science starts here.)
In 2026, three French gastroenterologists, Zlata Chkolnaia, Laurent Peyrin-Biroulet and Mathieu Uzzan, published a paper in the Journal of Crohn’s and Colitis titled “A descriptive comparison of phase 3 results and head-to-head trials in inflammatory bowel diseases.”
Inflammatory bowel disease (IBD) covers two chronic conditions, Crohn’s disease and ulcerative colitis, in which the immune system keeps attacking the gut. It often starts in young adulthood and can last a lifetime. Today there are several advanced drugs that work through different biological mechanisms. That’s a blessing for patients and a real headache for doctors deciding which one to use first.
The authors asked a very simple question: Does the subtraction shortcut actually work?
Their approach was elegant. At that point, only three head-to-head trials in IBD had compared two advanced drugs without a placebo arm. For each one, the authors went back to the original placebo-controlled trials of the same two drugs, did the subtraction, and noted what the shortcut would have predicted. Then they checked that prediction against what the head-to-head trial actually found.
In finance language, this is an out-of-sample test. You build your model on historical data and then check it against a period the model never saw.
And the results were humbling.
Test 1: Vedolizumab versus adalimumab in ulcerative colitis. The subtraction predicted that vedolizumab would beat adalimumab by about 17 percentage points in remission at one year. The actual head-to-head trial, called VARSITY, did find vedolizumab ahead, but by 8.8 points. Right direction, but the predicted gap was about double the real one.
Test 2: Ustekinumab versus adalimumab in Crohn’s disease (in patients who had never taken a biologic drug before). Here the subtraction predicted that adalimumab would win by 11.6 to 18.6 points, depending on which old adalimumab trial you used. The actual head-to-head trial, SEAVUE, found no statistically significant difference between the two drugs at one year. In fact, ustekinumab was slightly ahead numerically, by 4 points.
A small but important caution here. “No significant difference” does not mean the two drugs are proven identical. It only means the trial couldn’t confidently tell them apart. But it clearly does not support the big adalimumab advantage that the arithmetic promised.
Test 3: Risankizumab versus ustekinumab in Crohn’s disease (in patients whose earlier anti-TNF drug had failed). Here the shortcut could not produce any answer at all, because the subgroup data needed for the comparison had simply never been published. When the head-to-head trial, SEQUENCE, was finally run, risankizumab came out ahead by 15.6 points on endoscopic remission (healing seen when a camera examines the gut lining).
So the scorecard was one overestimate, one miss, and one blank!
To be fair, the authors are clear that their subtraction method was deliberately simple. It is not a formal statistical model, and they describe it as qualitative and illustrative. But they also looked at the more sophisticated method, called network meta-analysis, which pools many trials into one big statistical comparison. Those analyses didn’t agree with each other either. For vedolizumab versus adalimumab, published estimates of vedolizumab’s advantage ranged from 1.5 points all the way to 24 points. In Crohn’s disease, one major analysis favored adalimumab over ustekinumab, while another favored ustekinumab.
(If you’ve ever seen three sell-side analysts value the same company and land at three completely different target prices, this will sound very familiar!)
Why Did the Arithmetic Fail?
This is the most important section of the article, so stay with me.
The trials didn’t disagree because the drugs changed. They disagreed because the patients and the rules changed.
Take the ulcerative colitis trials. In the vedolizumab placebo trial, 48.2% of patients had already tried a biologic drug. In the adalimumab placebo trial, it was 40.3%. In the head-to-head trial, it was only 20.8%. Use of steroids at the start of the trial (a rough indicator of how sick and how heavily treated patients were) ranged from 36.1% to 77.9% across these trials. The trials enrolled patients across different years, from 2006 all the way to 2019, and standard care kept changing during that time.
Remember Fund A and Fund B, measured against different benchmarks in different market cycles? Exactly the same thing.
Then comes trial design, and this is where the finance analogy becomes almost uncomfortably accurate.
Many placebo-controlled trials use something called a re-randomization design. Everyone gets the drug for the first few weeks. Only those who respond move on to the long-term phase, where they are randomly assigned again, either to continue the drug or switch to placebo. So the one-year result you see reported is calculated only among patients who had already responded.
That’s survivorship bias, plain and simple. It’s like a fund database that quietly removes all the funds that shut down and then reports the average return of the survivors.

The authors show how big this effect can be. Suppose 40% to 60% of patients respond in the first few weeks, and 40% of those responders are in remission at one year. If you count everyone who started treatment, the real remission rate is only 16% to 24%. The head-to-head trials, on the other hand, mostly used a treat-through design, keeping every patient in the analysis from day one. Comparing those two numbers directly is like comparing one fund’s return on its winning positions with another fund’s return on its entire book.
There were other mismatches too. In the older Crohn’s disease trials, patients could take additional immune-suppressing medicines alongside the study drug. In the SEAVUE head-to-head trial, those medicines weren’t allowed.
Different patients, different eras, different rules. And then we subtract the headline numbers and act surprised when the answer doesn’t hold!
From “Which Drug Wins?” to “Which Drug for You?”
So far, the paper is about predicting one population-level result from another. That alone is a valuable warning. But it leads to a much bigger point, and I want to be honest that the next step is my own inference, not something the authors claim.
If clinical trial data can’t reliably tell us which of two drugs wins on average, it certainly can’t, on its own, tell us which drug will work best for one specific person.
To understand why, you need to know what a trial actually measures. Random assignment makes the two groups comparable. It does not make the individuals comparable. Each patient takes only one of the two drugs during the trial, so we never see what would have happened to that same person on the other one. What comes out is an average treatment effect: the difference between two group averages.
And averages can hide a lot. Back in 2004, Richard Kravitz and colleagues argued that a trial’s single headline benefit can hide a mixture of substantial benefit for some patients, little benefit for many, and harm for a few. They identified four reasons why people differ: their baseline risk of the bad outcome, how strongly their body responds to the drug, how vulnerable they are to side effects, and, interestingly, how much they personally value one outcome over another.
Let me make baseline risk concrete, because you deal with this every day without calling it that.
Imagine a drug that cuts the risk of a heart attack by 25% for everyone. For a high-risk person who starts with a 20% chance, that 25% reduction is worth 5 percentage points. You’d need to treat 20 such people to prevent one heart attack. For a low-risk person starting with a 2% chance, the same drug reduces risk by only half a percentage point. Now you’d need to treat 200 people to prevent one heart attack. Same drug, same “25%,” but a tenfold difference in real benefit, while side effects, costs and hassle stay about the same.

You already understand this. A 25% gain on $1,000 and a 25% gain on $100,000 are the same percentage but very different outcomes. Percentages without the base mean very little. Medicine works the same way.
The Trial Patient Is Probably Not You
There’s another problem: who gets into trials in the first place.
A 2020 systematic review by He, Morales and Guthrie looked at 305 trials across 31 medical conditions and estimated how many real-world patients would have been excluded by each trial’s entry criteria. The median was a whopping 77.1%. For asthma trials, it was 96%. And the most common reasons for exclusion were age, other illnesses, and taking other medications. In other words, the typical patient walking into a doctor’s clinic.
IBD is no exception. A 2012 study by Christina Ha and colleagues applied the entry criteria of major IBD drug trials to 206 real patients with moderate-to-severe disease. Only 31.1% would have qualified for any of them. Among Crohn’s disease patients who weren’t eligible for trials but received biologic drugs anyway, the response rate was 60%, compared with 89% among those who would have qualified. (That study was done at a single center by looking back at records, so it shows an association, not proof of cause. But the pattern is hard to ignore.)
It’s like backtesting a strategy only on large-cap, highly liquid stocks in calm markets and then selling it as suitable for every portfolio.
“Can’t We Just Look at Subgroups?”
You might ask: why not split the trial data into smaller groups and look for people similar to your patient?
Unfortunately, that’s a well-known trap. In a huge 1988 heart attack trial called ISIS-2, aspirin clearly saved lives. To make a point, the investigators split patients by astrological sign. Patients born under Gemini or Libra appeared to get no benefit from aspirin, while everyone else showed a striking benefit. Obviously, nobody believed stars affect how aspirin works. They included it to show how easily pure chance creates convincing-looking patterns.
Slice the data enough ways, and something will always look significant. In our world, we call it data mining.
So Are Clinical Trials Useless? Absolutely Not!
I want to be very clear here, because this is where people often get it wrong. None of this is an argument against clinical trials. Randomized trials remain the most reliable tool we have for finding out whether a treatment works. In fact, the Chkolnaia paper calls for more head-to-head trials, designed around comparisons that matter clinically and in populations that look like real patients.
The point is that trials answer one kind of question extremely well and another kind of question much less well.
For regulators, insurers and guideline committees (people making decisions for entire populations), trial averages are exactly the right tool. Policy should be built on averages.
But the doctor in the clinic isn’t setting policy. They are making one decision for one human being, with that person’s own disease history, earlier treatments, other conditions, appetite for risk, and priorities. Does this patient care most about getting better quickly? About avoiding injections? About a particular side effect? None of that appears in a trial’s headline number.
And this is not a rebellion against evidence-based medicine. It is evidence-based medicine. When David Sackett and colleagues defined the term in the BMJ in 1996, they described it as integrating individual clinical expertise with the best available external evidence. The evidence is where the decision begins, not where it ends.
What Did We Learn? Quick Summary
Based on all the evidence, here’s what I’d suggest you keep in mind:
- Treat trial results like a fund factsheet. They’re valuable, but always check who was included, over what period, and how the number was calculated.
- Ask whether you would have qualified for the trial. If people of your age, with your other conditions or your other medicines, were excluded, the result may not apply neatly to you.
- Always ask for absolute numbers, not just percentages. A “50% reduction in risk” means very different things depending on where your risk starts.
- Tell your doctor what matters to you. Speed, convenience, safety and cost are legitimate inputs into the decision, not distractions.
- Agree on a monitoring plan. For any one person, the most reliable evidence often comes from carefully tracking what happens after starting the treatment.
A Real Story to End With
Before I close, let me leave you with a real story.
Clopidogrel (sold as Plavix) is a widely used blood thinner given to reduce the risk of heart attacks and strokes. It went through large trials and showed clear benefit on average. But here’s the twist: clopidogrel doesn’t work as it is. The liver first has to convert it into its active form, mainly through an enzyme called CYP2C19.
Some people carry genetic variants that make this enzyme work poorly. They are called “poor metabolizers,” and an estimated 2% to 14% of the population falls into this group, with the rate varying by ethnic background. In these people, the drug simply doesn’t get activated properly, so they don’t get the full protection. In 2010, the US Food and Drug Administration added a boxed warning, its most serious type of warning, to the drug’s label. It advised doctors to consider other drugs or dosing strategies for patients identified as poor metabolizers.

Think about it. The drug looked fine on average in the trials. But for a specific subset of people, that average was misleading. And the only way to know whether you belonged to that subset was to look beyond the trial and at you.
What you should do is think about this story and try linking it with everything we’ve discussed. Trials tell us what works. A good doctor’s job is to find out what works for the patient sitting in front of them.
Your views and comments are welcome, and I’d be happy to respond.
