In Praise of Observational Evidence: Why RCTs Aren't Always the Gold Standard

While Randomized Control Trials are often hailed as the gold standard of science, I argue that observational evidence offers unique efficiency and power. From John Arbuthnot's 18th-century analysis of birth rates to James Lind's scurvy experiments, history shows we can uncover profound truths without expensive interventions. Observational studies often leverage massive sample sizes that modern trials cannot match, making them a superior tool for understanding complex phenomena in medicine and public health when RCTs are impractical.
It is a rare randomized trial that can detect a 0.3% difference in a binary random variable — but Laplace could, more than 200 years ago.
- derbOac
Experimental designs are critical for obvious reasons but they have a few critical flaws, that mostly all reduce to the fact you can't randomize manipulations with everything. Whether it be due to ethics or practical constraints, you can't conduct a RCT all the time.
This can be more subtly critical than it might seem, in that even if you can manipulate some proxy, often that proxy is insufficient in actually representing the phenomenon of interest, or the conditions under which they actually occur.
I often use the example of videogames and aggression. There were plenty of experimental studies of this but it was always questionable whether lab-induced anger is the same thing as, say, the sort of violence we generally are concerned about societally.
I generally have tried to teach students that experimental designs when done right provide powerful causal evidence of something, but often with limited generalizability; observational designs in contrast provide powerful generalizable evidence of some kind of association, but often with limited certainty about the causal pathways involved.
I've been in a department that was rabidly experimental in its focus and it always seemed sort of short-sighted, because people were idolizing RCTs with proxy manipulations that had questionable generalizability to the real-world phenomena they were trying to model.
Ideally you'd bring both experimental and observational evidence to bear on a question. Your conclusions should be robust to differ […]
- compiler-guy
One classic paper for on this topic:
_Parachute use to prevent death and major trauma related to gravitational challenge: systematic review of randomised controlled trials_
https://pmc.ncbi.nlm.nih.gov/articles/PMC300808/
"Advocates of evidence based medicine have criticised the adoption of interventions evaluated by using only observational data. We think that everyone might benefit if the most radical protagonists of evidence based medicine organised and participated in a double blind, randomised, placebo controlled, crossover trial of the parachute."
- mchusma
Good article. I will only add that I think the problems outlined in the article are exacerbated by the regulatory apparatus. If this were a debate related to truth seeking alone, it would be good. But trials drive banning/unbanning of treatments and medicines, as well as their mandated coverage by insurance.
We could be much more flexible in our approach to things if we would un-ban things. For example after phase 1 trial or based on sufficient observational evidence, things can no longer be banned but have a higher standard for "insurance is now required to cover it".
- khalic
Damn fine article, lovely conclusion, a real pleasure to read
- samuell
Quite thought-provoking, and connecting it to a related field it seems the (relative) success of LLMs and the likes are indications that enough data can at least learn you something without always needing to interfere with the world first(?)