Effect size expectations and common method bias

I think researchers in the social sciences often have unrealistic expectations about effect sizes. This has many causes, including publication bias (and selection bias more generally) and forking paths. Old news here.

Will Hobbs pointed me to his (PPNAS!) paper with Anthony Ong that highlights and examines another cause: common method bias.

Common method bias is the well-known (in some corners at least) phenomenon whereby specific common variance in variables measured through the same methods can produce bias. You can come up with many mechanisms for this. Variables measured in the same questionnaire can be correlated because of consistency motivations, the same tendency to give social desirably responses, similar uses of the similar scales, etc.

Many of these biases result in inflated correlations. Hobbs writes:

[U]nreasonable effect size priors is one of my main motivations for this line of work.

A lot of researchers seem to consider effect sizes meaningful only if they’re comparable to the huge observational correlations seen among subjective closed-ended survey items.

But often the quantities we really care about — or at least we are planning more ambitious field studies to estimate — are inherently going to be not measured in the same ways. We might assign a treatment and measure a survey outcome. We might measure a survey outcome, use this to target an intervention, and then look at outcomes in administrative data (e.g., income, health insurance data).

Here at this blog, perhaps there’s the most coverage of tiny studies, forking paths, and selection bias as causes of inflated effect size expectations. So this is a good reminder there are plenty of other causes, even with big samples or pre-registered analysis plans, like common method bias and confounding more generally.

This post is by Dean Eckles.

 

10 thoughts on “Effect size expectations and common method bias

  1. Interesting. I wonder whether some of this problem would be revealed by doing a simulation before collecting real data. The idea is that to construct the simulated data the first place, it is necessary to make some assumptions about nonsampling error, measurement error, etc.—and just about any reasonable assumptions will be better than the assumption, made implicitly by many methods, that these errors are zero.

    • Yes, I agree that would definitely be helpful. And that’s yet another example of why I like simulation for planning studies: it is so easy to add additional bits of realism in how the data is generated or will be analyzed.

      Then problems can still come later in convincing readers that the effect sizes you observe aren’t tiny, but that perhaps they have unrealistic expectations.

  2. Thanks for posting, I wasn’t aware of this. The survey example makes a lot of sense to me… you constrain the outcome space and so it’s not that surprising if individual differences end up creating correlations

  3. It is also important to point out that we are really interested in the “true” effect sizes and that these effect sizes are often attenuated considerably by method variance that is not shared across predictor and outcome variables. A single life-satisfaction item is surely not a 100% valid measure of life-satisfaction, let alone wellbeing. So, if we find that income correlates r = .1 with a single item life-satisfaction item, we are underestimating the real effect size a lot.

  4. The validation logic laid out in the Campbell & Fiske paper referenced by Fred Oswald is partly responsible for effect size inflation we see in personality/social psychology. If you’re clever, you’ creating your surveys and questionnaires in a manner that caters to humans’ desire to appear consistent and that exploits the stability of declarative memory. If you have a bit of an intuitive semantic sense of what items phrased in a specific manner will go with other, similar kinds of items, you can easily achieve high convergent validity coefficients, obtain strong, pure factors in factor analysis, and create highly consistent scales. And of course that is being exploited. Personality/social psychology journals therefore feature a plethora of studies where the now well-known sins that led to the replication crisis do not play much of a role. But the exploitation of common method bias and our left-hemisphere interpreter’s (nod to Gazzaniga) need to tell consistent stories does. This inflates effect size estimates based on the convergence of questionnaire scales grotesquely. Particularly when those questionnaires then fail to converge substantially with behavioral measures aiming to assess the same attribute (e.g., https://psycnet.apa.org/record/2013-35327-001; https://www.sciencedirect.com/science/article/abs/pii/S0272735811000997). The problem has been well known for a long time — to wit, McClelland’s 1972 paper provocativley titled “Opinions predict opnions. So what else is new?” (https://psycnet.apa.org/record/1972-31585-001). To wit, too, the Podsakoff et al paper about the dangers of common method bias cited by the authors of the PNAS paper. But it’s not like these problems are widely acknowledged in this quarter of psychology…

    • “Personality/social psychology journals therefore feature a plethora of studies where the now well-known sins that led to the replication crisis do not play much of a role. But the exploitation of common method bias and our left-hemisphere interpreter’s (nod to Gazzaniga) need to tell consistent stories does.”
      Good general point about the moving target of what problems are present in and biasing a research literature. There can be a whack-a-mole character to things.

    • It’s those pesky incentives again. Measurement and conceptual clarity is really, really hard in psychology, on account of (almost) nothing of theoretical interest being directly observable. But careful attention to measurement and design is not what pays off in publications.

      In psychological terms: Conceptual handwaving and high Cronbachs alpha in your sample data has been positively reinforced for a long time, and with intermittent reinforcement still in place the behaviour will presist.

    • Maybe you didn’t read Campbell and Fiske. They propose multi-method studies that remove the shared method bias.
      This approach has been used to show that correlations among Big Five scales are inflated by shared method variance.
      The approach also shows that implicit measures often lack convergent validity. Demonstrating convergent validity across distinct methods is a key step for validation that is often omitted. How do you control for method variance in your research?

Leave a Reply

Your email address will not be published. Required fields are marked *