A drug that reaches a pharmacy shelf has survived a filtering process designed, quite deliberately, to kill it. Understanding how clinical trials work means understanding that the failures are the point, and that the ones which succeed are the small residue left after most candidates have been thrown out.
A clinical trial is a controlled comparison in people. A group receiving the new treatment is compared against a group receiving something else — standard care, or a placebo — under rules fixed in advance, so that any difference between the groups can be attributed to the treatment rather than to luck, expectation or the way patients were chosen.
Everything else is machinery built to protect that comparison.
What a trial is actually testing
The question is almost never “does this work”. It is “does this work better than what we already do, in this specific group of people, measured this specific way”.
That is written down before anyone is enrolled, in a document called the protocol. It fixes who qualifies, how many people take part, how long the study runs, whether there is a control group, what dose is given, what will be measured and when, and how the results will be analysed.
Fixing it in advance is the safeguard. A researcher who can decide after the fact which outcome to report, or which patients to include, can extract an encouraging result from almost any dataset.
Before any of this begins with a new drug, the sponsor has to file an investigational new drug application with regulators, presenting the laboratory and animal data and the proposed protocol. Testing in humans is not something a company simply decides to start.
A second gate sits alongside the regulator. An independent ethics committee — an institutional review board in American usage — reviews the protocol and can refuse it, and every participant has to give informed consent after being told what is known, what is not, and that they may leave at any point without losing their normal care.
Underpinning both is a condition called equipoise: a trial is only defensible while the medical community genuinely does not know which arm is better. Once that uncertainty is gone, randomising patients means knowingly assigning some of them to the worse option.
The phases, and what each one buys
The sequence runs small to large, and each stage is answering a different question. The figures below are the FDA’s own.
| Phase | Participants | Length | Question | Proceed |
|---|---|---|---|---|
| 1 | 20–100, often healthy volunteers | Several months | Is it safe, and at what dose? | ~70% |
| 2 | Up to several hundred patients | Months to 2 years | Does it appear to do anything? | ~33% |
| 3 | 300–3,000 patients | 1 to 4 years | Is it better than standard care? | ~25–30% |
| 4 | Several thousand | After approval | What shows up in real use? | — |
Phase 1 is usually run in healthy volunteers, which sounds strange until you consider what is being asked. Nobody is testing whether the drug helps. They are establishing how much of it a human body tolerates and what happens as the dose climbs, and a healthy participant gives a cleaner read on that.
Cancer drugs are the standard exception, since giving a cytotoxic compound to a healthy person is not defensible. Those Phase 1 studies enrol patients who have the disease.
Phase 3 is where most of the safety data comes from, not Phase 1. Rare side effects are, by definition, invisible in a group of forty people; they only surface when thousands are exposed for years.
Multiply the phases together
Most coverage lists those phases and stops. The interesting number appears only when you combine them. Roughly 70% of drugs clear Phase 1. About a third of those clear Phase 2. Between a quarter and 30% of the survivors clear Phase 3.
Multiply the three and roughly six in every hundred drugs entering human testing come out the far end — around one in sixteen. That reframes the whole enterprise. A pharmaceutical pipeline is not a production line that occasionally jams; it is a sieve that discards almost everything, and the cost of the discarded candidates is carried by the few that survive.
It also explains where programmes die. Phase 2 is the first honest test of whether a compound does anything at all, and two thirds of them do not.
The attrition is not a sign the system is broken. A system that approved most of what it tested would be one that had stopped testing seriously.
Randomisation is the whole trick
If you let doctors choose which patients get the new treatment, they will choose — kindly and unconsciously — the patients most likely to do well. The comparison is then worthless, because the groups differed before the treatment began.
The fix is to remove the choice. In a randomised controlled trial, a computer assigns each participant to a group by chance alone.
The National Cancer Institute puts the logic plainly: randomisation, in which people are assigned to groups by chance alone, helps prevent bias. What makes it powerful is that it balances the things nobody thought to measure. Age and disease stage can be matched deliberately; the unrecorded differences in diet, genetics and stubbornness can only be balanced by chance.
This is why a large randomised trial outranks a much larger observational study. Numbers do not fix a selection problem. Only random assignment does.
Blinding handles what randomisation cannot
Knowing you are receiving a promising new drug changes how you report your symptoms. Knowing your patient is on the new drug changes how a doctor records their condition. Neither is dishonesty; both are entirely ordinary.
Blinding removes the knowledge. In a double-blind trial neither participant nor clinician knows who is in which group, which is why a placebo has to be indistinguishable from the real thing.
Blinding is not always possible. A surgical technique cannot be disguised, and a drug with an unmistakable side effect will often unblind itself. Trials in those situations are still worth running, but their results carry a caveat that rarely survives into the headline.
Because the participants and their doctors are blinded, somebody has to be able to see. An independent data monitoring committee reviews the unblinded results as they accumulate and can stop the trial early — for harm, or occasionally because the benefit is already so clear that continuing the control arm cannot be justified.
What is measured decides what can be claimed
A trial can only support conclusions about the thing it actually measured, and the choice of measurement is made by the people running it.
The distinction that matters is between an outcome patients care about — living longer, avoiding a heart attack, walking without pain — and a surrogate marker standing in for it, such as a blood test result or a tumour shrinking on a scan.
Surrogates are quicker and cheaper to measure, and sometimes they track the real outcome closely. Sometimes they do not, and a drug improves the number without helping anyone.
The same problem shapes any constructed statistic. Deciding what goes into the measurement determines what the resulting figure is capable of saying, which is exactly the issue underneath how inflation is measured.
So the useful question about a trial result is not how large the effect was. It is what was measured, and for how long.
Registration exists because of what used to happen
A trial that produces a disappointing result used to have a way of quietly not existing. It was run, it was filed away, and the published literature ended up as a biased sample of the research actually conducted — an effect known as publication bias.
The countermeasure is registration. Trials are recorded on a public registry such as ClinicalTrials.gov before they begin, with their design and intended outcome measures stated up front.
The register makes two things visible that were previously invisible: trials that were run and never reported, and trials whose reported outcome is not the outcome they set out to measure. Neither is detectable from the published paper alone.
Compliance is imperfect and enforcement is patchy, and researchers who study the problem still find substantial reporting gaps. The registry is a genuine improvement rather than a solved problem.
Approval is not the end of the testing
Phase 3 involves at most a few thousand people, monitored under trial conditions, usually excluding the very old, the pregnant and those with several conditions at once. Real use involves millions of people who look nothing like that.
Phase 4 is what happens afterwards: continued monitoring once a product is on the market, which is where genuinely rare harms and long-term effects are detected. Some of the most consequential drug findings in history have come at this stage, years after approval.
It is also the stage most likely to be missing from a news story, because the interesting moment for coverage is the approval itself.
The evidence base for any treatment keeps moving after the licence is granted, which is why standard advice shifts — as it has repeatedly for what the evidence on sleep actually supports.
What to look for when a trial makes the news
Coverage compresses all of the above into a single effect size, usually the most flattering one available. Four questions restore most of what was lost:
- What was the comparison group given — a placebo, or the best existing treatment? Beating nothing is a much weaker claim
- Was the reported outcome a real endpoint or a surrogate marker
- How many participants, and for how long, relative to the condition being treated
- Is the effect reported in absolute terms, or only as a relative reduction that sounds larger than it is
And one structural check that costs nothing: a Phase 2 result is a reason to run a Phase 3 trial, not a reason to change what anyone does. Two thirds of the compounds that look promising at that point will not survive the next stage.