As a statistician, I was trained to think of randomized experimentation as representing the gold standard of knowledge in the social sciences, and, despite having seen occasional arguments to the contrary, I still hold that view, expressed pithily by Box, Hunter, and Hunter that “To find out what happens when you change something, it is necessary to change it.”
At the same time, in my capacity as a social scientist, I’ve published many applied research papers, almost none of which have used experimental data.
In the present article, I’ll address the following questions:
1. Why do I agree with the consensus characterization of randomized experimentation as a gold standard?
2. Given point 1 above, why does almost all my research use observational data?
. . .
For the rest, you can read the article, written in question is from 2010, ultimately published in a book in 2014, and still relevant today, I think! For more on this perspective you can see chapters 18-21, the causal inference chapters of Regression and Other Stories.
Also relevant is the idea that mathematical and statistical reasoning are themselves experimental, as discussed in this article on Lakatos and in this talk, “When You do Applied Statistics, You’re Acting Like a Scientist. Why Does this matter?”
Randomized experiments without the context of some model of whats going on aren’t very useful, and can even be misleading.
My example is that (pretty much?) all cancer treatments that appear somewhat effective have the “side effects” of reduced appetite, nausea, and vomitting. This is regardless of the supposed mechanism of action. Then there are about a million papers on how caloric restriction slows tumor growth. Along with cellular/molecular observations known for 100 years like the Warburg effect.
So why is there not a single trial that monitors nutrient intake/absorption during treatment to see if that could be the mechanism? Maybe it is only partly the explanation. But if you can halve the dosage (along with restricted diet) that could yield huge benefits in terms of money and toxicity.
Instead we use the EBM (“brute force”) approach that applies the RCT algorithm to each new thing without considering prior information.
I think that the practical obstacles are seen as insurmountable. Sure, there is a ton of randomized animal data and observational human data on the benefits of calorie restriction. However, controlling calorie intake in humans is hard. Dietary interventions have at best so-so results in actually reducing dietary intake in real people. To do a trial, you would have to lock the patients up with severe restrictions on visitors to ensure being able to limit intake. I think that the families of people with lung cancer would certainly try to smuggle in some chocolate bars for a person with two kilos of weight loss. Parenthetically, my wife keeps buying chocolate despite my entreaties.
Good clinical trials are hard to do. That is why statistical vigilance is so important.
Nope. Many dont have those side effects. Pembrolizumab, trastuzumab, exemestane, leuprolide, etc etc. Your premise is trivially wrong.
> “To find out what happens when you change something, it is necessary to change it.”
I don’t see how this statement from BHH can be taken as an expression of support of *randomized* experimentation.
This seems like a pretty obvious basic requirement of an experiment to estimate effects whether you believe ‘randomization’ is a necessary property or not.
Jyd:
I agree that the important part is the experimentation, not the randomization. You could just remove the word “randomized” from the first sentence of the above post.
> “To find out what happens when you change something, it is necessary to change it.”
> This seems like a pretty obvious basic requirement of an experiment to estimate effects whether you believe ‘randomization’ is a necessary property or not.
To estimate effects it’s also necessary to know what would have happened if you had not changed that thing. Finding out what happened when you did is not enough.
Interesting. I’m tempted to write my own version:
1. Why do I agree with the consensus characterization of randomized experimentation as a gold standard?
2. Given point 1 above which has led me to do a lot of experiments in my career, why do I roll my eyes when people present me with new randomized experiments or ideas for new randomized experiments?
Jessica:
I think one clue is that we also roll our eyes when people present us with new analyses using regression discontinuity, instrumental variables, difference-in-difference, path analysis, synthetic controls, etc. Tools for causal identification are often used as an excuse for researchers to turn off their brains and avoid thinking hard about the connections between their quantitative claims and the purported causal mechanisms.
I’m reminded of the well-known statistical issue that, when a decision is made based on two factors, there can be a negative correlation between those factors in a selected dataset. For example, on a basketball team you could see a negative correlation between height and shooting ability, or in college students you could see a negative correlation between high school grade point average and admissions test scores, because of selection in who is accepted and decides to attend. If a study is strong on identification (for example, with a randomized design), this allows it to be weak in other places. Of course a study can be weak in all dimensions, but then you’re less likely to have heard about it in the first place; that’s the selection.
Whenever Imre Lakatos’s name comes up, I think it is necessary to point to this unsettling commentary:
https://www.dropbox.com/s/2sivwbs683gc3nz/ROH-Lakatos.pdf?dl=0
“All in all, a lot of the evidence points to the possibility that Lakatos was a psychopath, which is indeed how he was described by Dr. Klára Majerszky, who worked at the National Psychiatric and Neurological Institute in Budapest and who knew Lakatos person-ally before the war.”
“Lakatos, who was then already a lecturer at the London School of Economics, was still collaborating with the Hungarian secret police”
“Lakatos’s real-life actions that caused loss of life and immense human suffering have barely produced a yawn among philosophers”
Paul:
Yes, we discussed this on the blog a couple years ago.
That is interesting, I don’t follow your point though. Why do you find it necessary?
Do you realize the history of statistics is deeply intertwined with eugenics? To me such things have little to do with whether an idea/argument is correct or interesting.