“Published estimates of group differences in multisensory integration are inflated”

Mike Beauchamp sends in the above picture of Buster (“so-named by my son because we adopted him as a stray kitten run over by a car and ‘all busted up'”) sends along this article (coauthored with John F. Magnotti) “examining how the usual suspects (small n, forking paths, etc.) had led our little sub-field of psychology/neuroscience, multisensory integration, astray.” The article begins:

A common measure of multisensory integration is the McGurk effect, an illusion in which incongruent auditory and visual speech are integrated to produce an entirely different percept. Published studies report that participants who differ in age, gender, culture, native language, or traits related to neurological or psychiatric disorders also differ in their susceptibility to the McGurk effect. These group-level differences are used as evidence for fundamental alterations in sensory processing between populations. Using empirical data and statistical simulations tested under a range of conditions, we show that published estimates of group differences in the McGurk effect are inflated when only statistically significant (p < 0.05) results are published [emphasis added]. With a sample size typical of published studies, a group difference of 10% would be reported as 31%. As a consequence of this inflation, follow-up studies often fail to replicate published reports of large between-group differences. Inaccurate estimates of effect sizes and replication failures are especially problematic in studies of clinical populations involving expensive and time-consuming interventions, such as training paradigms to improve sensory processing. Reducing effect size inflation and increasing replicability requires increasing the number of participants by an order of magnitude compared with current practice.

Type M error!

4 thoughts on ““Published estimates of group differences in multisensory integration are inflated”

  1. Take it as a principle that there is always a (possibly negligible) difference between groups.

    Increasing the sample size an order of magnitude may end up in “too much” significance. If not 10x, it could be 100x, or 1000x. There is a point where everything becomes significant.

    There is a balance between the typical magnitude, variability and measurement precision in a field vs sample size and significance cutoff. Substantially increasing only the sample size will throw off this balance. Eg, instead of ~1/3 of studies resulting in a discovery, with sufficent n then nearly every experiment becomes significant. Historically, the cutoff will change from 0.05 to 0.01, or whatever. Then you are back where you started with the significance filter.

    Like Jupiter’s moons, there was great excitement when the first four were seen with a blurry telescope but few people cared when the 80th moon was recently discovered (and it has been estimated there are hundreds more smaller moons). Likely the cutoff for what counts on a moon will be made more stringent in the future.

    • no… just no no no

      The point of the increased N is to get a more accurate estimate of the actual effect. There is no trade off of too much statistical significance. In addition to moving toward a literature that doesn’t have inflated effects we’re also supposed to be moving toward a literature where statistical significance counts so much that a completely trivial effect is ignored. Therefore, no, you are not back to where you started with a significance filter. If every study is powered to near 1.0 that is not the same as having .30 power and only publishing significant findings.

      That said, the bigger issue here is that the paper should also be highlighting that publishing non-significant findings is important in understanding the effect size to a greater degree. Every bit of data from a well designed study contributes to the estimate and, given that they are usually publicly funded, results in information that must be publicly available..

      • The point of the increased N is to get a more accurate estimate of the actual effect. There is no trade off of too much statistical significance.

        There shouldn’t be, but there is. If it is too easy to make a “discovery”, people will start sensing something is wrong (there is, it is testing a strawman hypothesis). Then the significance threshold will be lowered and this new level will be used to filter the results. This is the actual behavior we have observed in practice.

        we’re also supposed to be moving toward a literature where statistical significance counts so much that a completely trivial effect is ignored

        How does this work?

Leave a Reply

Your email address will not be published. Required fields are marked *