Joe Simmons, Leif Nelson, and Uri Simonsohn agree with us regarding the much publicized but implausible and unsubstantiated claims of huge effects from nudge interventions

We wrote about this last year in our post, PNAS GIGO QRP WTF: This meta-analysis of nudge experiments is approaching the platonic ideal of junk science and our followup PNAS article, No reason to expect large and consistent effects of nudge interventions:

The article in question is called “The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains” . . . From the abstract of the paper:

Our results show that choice architecture interventions [“nudging”] overall promote behavior change with a small to medium effect size of Cohen’s d = 0.45 . . .

Wha . . .? An effect size of 0.45 is not “small to medium”; it’s huge. Huge as in implausible that these little interventions would shift people, on average, by half a standard deviation. I mean, sure, if the data really show this, then it would be notable—it would be big news—because it would be a huge effect.

The claim of an average effect of 0.45 standard deviations does not, by itself, make the article’s conclusions wrong—it’s always possible that such large effects exist—but it’s a bad sign, and labeling it as “small to medium” points to a misconception that reminds us of the process whereby these upwardly biased estimates get published.

The article in question referred to about 200 previous papers, including 11 articles authored or coauthored by the discredited food researcher Brian Wansink, and also a notorious retracted paper coauthored by the controversial dishonesty researcher Dan Ariely.

But, I continued last year:

I would not believe the results of this meta-analysis even if it did not include any of the above 12 papers, as I don’t see any good reason to trust the individual studies that went into the meta-analysis. It’s a whole literature of noisy data, small sample sizes, and selection on statistical significance, hence massive overestimates of effect sizes. This is not a secret: look at the papers in question and you will see, over and over again, that they’re selecting what to report based on whether the p-value is less than 0.05. The problem here is not the p-value—I’d have a similar issue if they were to select on whether the Bayes factor is greater then 3, for example—; rather, the problem is the selection, which induces noise (through the reduction of continuous data to a binary summary) and bias (by not allowing small effects to be reported at all).

tl;dr. It’s a literature full of junk, and the inclusion of 12 discredited papers is a problem, not just in itself, but in that it indicates the lack of care or quality control that went into putting together the papers that went into the meta-analysis. Either the authors carefully vetted each paper in the meta-analysis—in which case, they did a rotten job of vetting, given the Wansink and Ariely papers that got in—or they didn’t vet, in which case I don’t know why we’re supposed to believe the others there. So, no, I don’t believe any of the conclusions of that meta-analysis. GIGO.

Still and all,

I indeed think that “nudging” has been oversold, but the underlying idea—“choice architecture” or whatever you want to call it—is important. Defaults can make a big difference sometimes.

Here’s how we put it in our PNAS article:

Nudge interventions may work, under certain conditions, but their effectiveness can vary to a great degree, and the conditions under which they work are barely identified in the literature. . . . the authors [of the meta-analysis under discussion] focus their conclusions on this average value and on subgroups, leaving aside the large degree of unexplained heterogeneity in apparent effects across published studies. For example, despite the analyses above being consistent with a large proportion of studies having near-zero underlying effects, the authors conclude that nudges work “across a wide range of behavioral domains, population segments, and geographical locations.”

We summarized:

As a scientific field, instead of focusing on average effects, we need to understand when and where some nudges have huge positive effects and why others are not able to repeat those successes. Until then, with a few exceptions [e.g., defaults], we see no reason to expect large and consistent effects when designing nudge experiments or running interventions.

New post by Simmons, Nelson, and Simonsohn

More recently, the psychology researchers Joe Simmons, Leif Nelson, and Uri Simonsohn wrote a post agreeing with us! They refer to the meta-analysis as estimating a “meaningless mean” (a point which I agree with regarding nudges, for reasons discussed above, and also more generally, as discussed here).

In their post, they link to our PNAS article and write, “letters published in PNAS, responding to this article, have proposed that the overall average effect may be much smaller – perhaps as low as zero – when correcting for publication bias.” Just to be clear, in our article, we do not propose an “overall average effect” of nudging; rather, as indicated in the above-quoted passage, we argue against the idea of talking about an “average effect” and we fault the article under discussion for attempting to use their meta-analysis to make such broad claims.

Simmons et al. go beyond what we did by looking in detail at a few of the papers cited in the meta-analysis. They find that it does not make much sense to combine the effect sizes from these different studies. I’m not surprised—see my comments above about vetting—but there’s a big difference between “I’m not surprised” and actually seeing the details, so I think Simmons et al. have made a useful contribution by getting down and dirty and reading those papers.

What they found does not change our conclusions—as Simmons et al. say in footnote 10 of their post, their conclusions are consistent with the point we make in our PNAS article that “instead of focusing on average effects, we need to understand when and where some nudges have huge positive effects and why others are not able to repeat those successes”—but, again, it’s good to see this common-sense conclusion backed up by a careful look at some of those individual studies.

You might ask, Why didn’t the authors of the original meta-analysis look so carefully at the individual papers they were citing? A quick answer is that, had they done so, they wouldn’t have been able to make those dramatic claims that got their paper published in PNAS. A longer answer is selection bias: the kinds of researchers who read the individual papers carefully won’t be publishing articles making those sorts of ridiculous claims, and unfortunately in the current academic and science-communication media environment, it’s often the outrageous claims that get the attention. Not always—skepticism can sometimes get publicity too—but I think the balance still favors extreme and implausible claims over reasoned common sense.

So, yeah, it’s good to see Simmons et al. putting in the effort to read those papers. You might say that this was way too much effort to spend following up on an piece of GIGO, but recall the Javert paradox and recall what they say about the dead horse. Sometimes it’s valuable for researchers to put in their expertise to explain how other people can get things so wrong. I’m a statistician so I focused on the statistical problems of working with a bunch of estimates with large and unknown biases, estimating different things; Simmons et al. are psychologists so they’re focusing on the substance of how those studies are different. And we’re in agreement that it’s annoying when researchers seem to think that statistical tools such as meta-analysis can resolve major data problems or major conceptual problems. Again, GIGO.

12 thoughts on “Joe Simmons, Leif Nelson, and Uri Simonsohn agree with us regarding the much publicized but implausible and unsubstantiated claims of huge effects from nudge interventions

  1. What you wouldn’t do is consult an average of a bunch of different effects involving a bunch of different manipulations and a bunch of different measures and a bunch of a different contexts. And why not? Because that average is meaningless.

    Same thing is done in the “multiverse” analyses. You can multiple models with a “B1” coefficient that weights the same x1 variable:

    y = B1*x1
    y = B1*x1 + B2*x2
    y = B1*x1 + B2*x2 + B3*x3
    y = B1*x1 + B3*x3
    […]

    Then there can be interactions, non-linear models, and so on all with a B1 coefficient. What is the meaning of the average B1?

    Perhaps say the data was generated by something like newtons law of gravitation: y = (B1*x1*x2)/x3^2

    Would the correct model even be included in the set being averaged?

    • I suppose that if you used the multi-verse analysis to derive an average B1 you’re absolutely correct. I never got the impression that’s what it is for though. Perhaps I read wrong but my impression was that it was to see the distribution of p-values to circumvent an accusation that a selcted p-value used was cherry picked and also to just see the potential effect of researcher choices.

      Of course, including an interaction with B1 and using the model coefficient significance would be ludicrous to include in either event.

  2. I’d still be shocked if there wasn’t a fucking huge effect of changing thing like organ donation to an opt-out system (especially if you had to raise the issue to opt-out).

    I suspect that might not be a nudge in the sense you are concerned about but I think some people understand a nudge as something that’s changing behavior in a way that’s no intuitively coercive. I mean it might be annoying to have to remember to object to organ donation but you aren’t intuitively being coerced.

  3. a central point is the following of blind technical criteria.

    researchers try to follow technical rules also to avoid subjective judgement.

    it’s a challenging trade off at times

    and saying “I’m doing it technically except this wansink papers” adds a level of choice. should we do bonferonni? most choices are framed as obvious.

    nothing is easy. but important to note.

    • Jazi:

      Indeed, it’s difficult. As discussed in my earlier post and in more detail by Simmons et al., the problem is not just with the Wansink and Ariely papers. It’s the whole published literature, or a large chunk of it, that has huge overestimates. Again, this is not just a problem in popular books such as Nudge, Mindless Eating, and the collected works of the Sage of Durham, but also in published journal articles.

      Given that population of published studies, just about any meta-analysis that blindly uses published results will be hopeless.

  4. I hope that psycholinguists are also reading this post and are able to extrapolate to their own field. We have exactly the same type of problem in psycholinguistics: exclusive focus on a supposed but entirely fictional average effect, and then theory building based on that imagined effect existing “on average”. The whole field is built on this misunderstanding about what the target of modeling is.

    How to proceed if the average doesn’t really cut it? One thing my former student Himanshu Yadav and I tried to do recently is to take the published studies on a particular topic and try to figure out which of the different competing computational models can explain the variation best across studies. I thought that this was a better way to attack the question of “does the effect exist”, by looking at whether models (given some regularizing priors) can explain the data better than other models.

    This is still not very satisfying because most (all) of the published studies (some are mine) are just GIGO, with noisy overestimates that really should not be the target of modeling. Luckily a few people are now starting to obtain less noisy estimates, by either increasing sample size or making the manipulation stronger or both, but in my own lab many of the most famous effects do not replicate and there is even evidence against them (some really damaging papers will become public soon, damaging for existing psycholinguistic theories’ overblown claims). But until higher precision studies become the norm, this was the best we could manage.

    Here’s the paper: https://www.sciencedirect.com/science/article/pii/S0749596X22000870?via%3Dihub

    Another interesting thing to do instead of thinking about an imaginary average effect is just to try to model the individual level effects (the shrunken estimates of individual level effects from a hierarchical model). This seems like a really informative and constructive way to go, but one needs a computational process model, hand-waving models (which are the norm in psycholinguistics) just won’t cut it.

    Here is that paper: https://direct.mit.edu/opmi/article/doi/10.1162/opmi_a_00052/110716/Individual-Differences-in-Cue-Weighting-in

    I decided to publish it in Open Mind despite the journal’s ridiculous title because I thought the paper was way too ahead of its time and people were not ready for it, given their obsession with modeling average effects (not even taking the uncertainty of that effect estimate into account).

    PS I still publish in Elsevier journals sometimes, as in the first case above, just to show that I can if I want to; funding agencies apparently only take you seriously if you can demonstrate that you can publish in major commerical journals. There is a new serious journal, Glossa Psycholinguistics, that is probably the future for our field.

    PPS Surprisingly, Open Mind, despite being an MIT Press journal, did a really bad and sloppy job with the proofing of the paper. Also, for the longest time they had the proofs up *on the journal page*, with draft written across each page. I had to contact the editor in chief to get this fixed. Usually that level of incompetence is only something that Elsevier and Routledge etc. exhibit. So open access journals still have a way to go.

  5. Another literature bears similarity with the nudge literature in these respects: lock-in. Studies of markets with network externalities (e.g., social media, telecommunications, the QWERTY keyboard, beta/VHS standards wars, etc.) commonly cited “lock-in” as a market failure: once a standard becomes widely adopted, it is difficult to replace even if a superior product/service becomes available. An economist, Liebowitz (https://scholar.google.com/citations?user=TXJIckYAAAAJ&hl=en), reexamined those claims, distinguishing between “strong” and “weak” lock in. While there was little empirical content, the idea is that lock-in occurs more strongly when the stakes are not great – the larger the stakes, the harder it is for people to get locked in to an inferior standard. Similarly, I suspect that nudges work most strongly when the stakes are low – the larger the consequences, the smaller the effectiveness of simple nudges. I’m sure there are exceptions, but this makes the average effect irrelevant. I think the expectation should be that nudges will exhibit stronger effects when they don’t matter too much. Of course, it is much easier to get attention when citing large effects of nudges or lock-in for every circumstance in which they may be present.

  6. Andrew today mentions

    “Simmons, Nelson, Simonsohn”

    They, it turns out, are in the thick of uncovering the current Harvard University problem regarding data manipulation. Because of this, I have been reading the book, “Rebel Talent”, by Francesca Gino, the person most under fire. Putting aside any issue of falsification of data, the book is very well written and engrossing. Nevertheless, she seems to be a true believer when it comes to: separate the subjects (somehow, but not necessarily at random) into two groups, treatment and non-treatment, respectively, and

    low p-value implies “significance”

    which is not just applicable to Harvard undergraduates but to the universe as a whole.

    One of her examples (page 16) has to do with wearing red (Converse) sneakers in one section of a Harvard class and more traditional foot gear in the preceding class. Her conclusion from this successful experiment is “…we all share the desire to be happy…And something as simple as a pair of red sneakers might make all the difference.”

Leave a Reply

Your email address will not be published. Required fields are marked *