Why quantitative understanding of effect sizes matters, even if all you care about is the presence of the effect

In reaction to my article with Andy King proposing post-publication review, Dan “Fast and Frugal” Goldstein writes:

Your process limits information search, computation, and time so it seems fast and frugal to me. Happy you still associate me with that term. It was something I coined as a grad student.

In your proposal, only hit papers get audited. It reminds me a bit of Mel Brooks’ The Producers in which the protagonists use the logic “who would audit a flop?” and stay under the radar by intentionally producing a bad show. Fraudsters have likely attempted to make their work seem worthy of publication while ensuring it doesn’t attract too much attention. Just like in The Producers, though, this sometimes backfires.

I don’t know about that! My impression with fraudsters is that they think that fraud is normal science, perhaps out of some mixture of bad education in research methods, a view that “everybody does it,” and a general lack of understanding of how non-cheaters (like you and me!) think. Think of people like Wansink who gave general advice to to the world on how to p-hack, or Gino and Ariely, who published papers on dishonesty, or Mary Rosh, who surely believes that whatever shady statistical manipulations she does are nothing compared to the dastardly deeds done by the Democrats.

I’m sure there’s tons of below-the-radar cheating and bad science that we don’t hear about, but a fair number of prominent science fraudsters seem to enjoy the limelight. One reason for this seemingly self-sabotaging behavior, I think, is that cheating enabled these people to attain great professional success for years. They had no reason to think the juice would stop flowing.

To return to my proposal with Andy King: I think it’s ok that only the hit papers get audited. Bad papers that get no intention aren’t doing much damage, right?

Goldstein adds:

By the way, I was just having a conversation about your sensing that something was amiss with the LaCour study. For years now I have used this quote of yours in a talk I give about putting numbers into perspective. I argue that it’s really important that people learn how to put numbers into perspective because if they don’t, they won’t notice that something is unusual and worthy of a deeper audit. You somehow sensed something was up with the Lacour result. You didn’t think it was fraud yet but you knew it was strange because you know how to put such differences into perspective:

A difference of 0.8 on a five-point scale . . . wow! You rarely see this sort of thing. Just do the math. On a 1-5 scale, the maximum theoretically possible change would be 4. But, considering that lots of people are already at “4” or “5” on the scale, it’s hard to imagine an average change of more than 2. And that would be massive. So we’re talking about a causal effect that’s a full 40% of what is pretty much the maximum change imaginable. Wow, indeed. And, judging by the small standard errors (again, see the graphs above), these effects are real, not obtained by capitalizing on chance or the statistical significance filter or anything like that.

My colleagues and I recently wrote a paper on this general topic of average effect sizes. It’s our contention that people generally are way too optimistic about possible effect sizes, in large part because they don’t think about variation. If you ask someone to hypothesize an effect size, you’ll typically get a guess of the largest effect that might occur.

But what if you don’t really care about effect size–you just want to know about the effect?

For example, maybe you don’t believe that women during certain times of the month are three times more likely to wear red or pink shirts, but you are interested in some sort of evolutionary psychology theory of sexual display. In that case, why should the effect size matter? Why care that a study reported an estimate that was ridiculously implausible?

I have two to this questions, and thus two reasons why effect size is important even for problems where you don’t directly care about effect sizes:

1. Effect sizes vary. An treatment that has an effect (that is, a true effect, not just an estimated effect) of 0.1 for one group of people in one setting could have an effect of -0.2 in some other scenario. A treatment effect in an experiment is the sum of all sorts of things, positive and negative, and there’s no logical reason to think the sign of the effect will be preserved. Effect size matters. The issue is not just that a smaller and more realistic effect size is less important; it’s also that smaller effects can be more easily produced by other factors, and this reduces the generality of any claims, even if the experiment at hand was done well.

2. Experiments produces a standard errors as well as estimates. If the standard error from a study is large compared to any realistic effect size, then the study contains very little information. Effect size is important in understanding the informativeness of an experiment, and to do this right you need to have some sense of what the true effect size could be. You can’t just use a point estimate from the study itself, as this estimate will inherently be too noisy to use to judge the information in the study. As I wrote in this note for the Annals of Surgery, Post-hoc power using observed estimate of effect size is too noisy to be useful.

7 thoughts on “Why quantitative understanding of effect sizes matters, even if all you care about is the presence of the effect

  1. Ignoring effect size is indefensible. Suppose you are “only” interested in whether the effect exists – e.g., does news about shark attacks influence election results? If you ignore effect size to determine this, you are accepting the idea that any effect – or a statistically significant effect, if you are at least a little bit careful – is evidence that the effect is real. But as you say in #2, given that there is variation, a small effect size can easily mask a spurious result. You have to rely on statistical significance as a defense that the effect is real, but most of us reject that standard. Given the garden of forking paths, a statistically significant but small effect is inadequate for the claims that researchers might want to make. At best, it is an invitation to conduct further research. At worst, it is an excuse for ignoring all the assumptions that were needed to produce the small effect that was found.

    • I don’t disagree with the points you and Andrew are making (FWIW), but caution is advised about blanket statements. I’m thinking of the research question posed by an actual undergraduate in a social science class, “does viewing porn make men more likely to commit rape?” To me, this is the perfect question to “prove” (test) the “effect size matters” rule. If a net positive result means that a single excess rape occurred due to porn, that is already one too many excess rapes. Heck, I might not even need statistical significance to make an argument for banning porn based solely on the precautionary principle and a lack of any serious arguments about the benefits of porn.

      • And if a net negative result means that a single excess rape was prevented due to porn, would you encourage subsidizing porn? I’m only being partly facetious because I think your example is a bit unhinged.

      • Matt
        Yes, I find your example truly bizarre. It is not hard to get an empirical study to show that X might cause a statistically significant increase of 1 additional truly horrible act (rape, murder, mass murder, scamming the elderly out of their life savings, etc.). Given the Piranha principal (perhaps not the best example or the correct principal to cite), there are probably dozens of other things Y, Z, … that would cause -1 of these horrible actions. Isolating viewing porn because it is associated with one additional rape ignores the fact that viewing porn might substitute for any number of car rides and an expected death toll from increased accidents that exceeds 1. Or Will’s possibility that instead of a +1 effect, it is -1.

        The real problem with your example is that thinking like that leads to all kinds of absurd conclusions. If we find one case where an e-bike leads to an irresponsible riding and a death, then we ban e-bikes. If one person dies due to contaminated food, we shut down the business that delivered that food. It is the antithesis of statistical thinking – find an example and the conclusion follows. Note that your example concerned the effect size of 1 rape, based on some statistical study. My example can rest on one observed case with a known cause – no analysis is required since we have the concrete case.

  2. Seconding Dale that ignoring effect size is indefensible, but for a different reason: you want to make falsifiable predictions. To make your prediction more falsifiable, you need to predict the magnitude of an effect, and ideally the magnitude as whatever control parameter you have varies (that is, you want to predict an entire curve). This is the best way to learn from your data, make as sharp of a prediction as possible so you get a hint as to how to fix it when it fails.

    Not only is the earth round, but it has a radius of about 6000 km. Not only is the earth tilted, the tilt is about 23 degrees. Etc.

    • Most undergrad econometrics courses go out of their way to make the distinction between statistical significance and economic significance. In my experience, a lot of students don’t really pick up on the difference.

      “The statistical significance suggests that the coefficient is non-zero, but the coefficient is so small that, really, who gives a shit? It isn’t economically significant, even if there’s some reason to believe it is non-zero.”

      “Yeah, but then why do we say it is statistically significant?”

      It’s like the word “significant” dominates the distinction. And I gather some researchers like to exploit that.

Leave a Reply

Your email address will not be published. Required fields are marked *