“I think there’s an argument to be made that much meta-scientific work is a kind of mirror image of the empirical work it critiques”

In the context of our recent discussion of the p-curve paper, Richard Morey wrote, “I think there’s an argument to be made that much meta-scientific work is a kind of mirror image of the empirical work it critiques,” and he shared this chart:

I think Morey is on to something here, but, as someone who does a lot of empirical science and a lot of meta-science, I think there’s one big thing he’s missing, one major asymmetry between empirical science and meta-science, and that is that bad empirical science makes strong claims, and the role of meta-science is to question the evidential support behind these claims, not usually to make a positive claim in itself.

The usual pattern goes like this: empirical scientists collect data D, perform analysis A, and use these to make strong general claim X about the world. The meta-scientist then comes along to assess the evidence. A negative meta-science analysis comes to the conclusion that D + A do not provide good evidence for X. The meta-science analysis does not make the strong claim that X is false, let alone the even stronger claim that some preferred alternative Y is true.

This comes up all the time. Some Cornell psychology professor claims to have strong evidence for extra-sensory perception or influence of food labeling on eating or whatever. The meta-scientist comes along and notes irregularities with the data or analysis and provides an alternative story of how these apparently convincing patterns in data could have come to be. The conclusion of the meta-scientific report is not that ESP or large effects of food labels don’t exist but rather that the published record does not provide good evidence of these extraordinary claims. (And indeed the claims are extraordinary, which is how they got so much publicity in the first place.)

It’s the all-important distinction between truth and evidence. I know that Morey understands this distinction and I’m not saying that anything in his above chart is wrong; I’m just trying to put it in the larger perspective of scientific inquiry.

In discussing the above asymmetry between empirical science and meta-science, I’m not saying that meta-science is better. Meta-science is fundamentally parasitic on empirical science, and, sure, empirical science is associated with bold claims, but it’s through making bold leaps–and being willing to retract those leaps as needed–that we make progress. The problem with bad science is not so much the overconfident conjectures–such steps may be psychologically necessary–so much as the unwillingness to reflect on contrary evidence, the unwillingness to admit error, and the practice of not confronting past mistakes.

And also the really stupid things that people say and never apologize for.

15 thoughts on ““I think there’s an argument to be made that much meta-scientific work is a kind of mirror image of the empirical work it critiques”

  1. Quote from the blog post: “I think there’s an argument to be made that much meta-scientific work is a kind of mirror image of the empirical work it critiques,”

    Ah, that’s reassuring to me to read in the sense that I may not be the only one to think about possible problematic issues concerning “meta-science”. At a certain point in time I had come across multiple things that seem to me to possibly be “conceptual replications” of problematic issues. Some of these were proposed as being “solutions” or “improvements” in light of the so-called “replication crisis”. I thought it might be useful and/or amusing to make a manuscript about all of these things, in which I also ask many questions. Some of these questions might resonate with the topic of this blog post:

    “Is there a risk that “meta-science” and/or evaluation research can be used (in the short-term) to “nudge” or “steer” (parts of) Psychological Science in a certain direction without much and/or sufficient thought concerning the soundness, desirability, and/or validity (in the long-term) of the object or process under investigation? Is there a risk that the prospect of future “meta-scientific” and/or evaluation research can be used to ignore, brush aside, and/or neglect (possibly and/or probably) valid criticism? Can “meta-scientific” and/or evaluation research have similar problems as “normal” research concerning its validity, conflicts of interest, possible corruption, soundness of design and conclusions, etc.?”

    • Anon:

      Yes, one thing I’ve discussed in the past is the way that the meta-science movement in psychology has been an awkward coalition of reformers who wanted to use meta-science to discredit a bunch of old research, and status-quo types who wanted to use meta-science as a way to bolster the credibility of that old work. Both groups were reacting to the replication crisis but in different ways, and I think it’s been difficult for Brian Nosek and others to keep the support of both factions.

      • Some of these “reformers” might provide great examples concerning the quote “I think there’s an argument to be made that much meta-scientific work is a kind of mirror image of the empirical work it critiques,”

        If I am not mistaken, there was a blog post by Bastian in 2017 for instance about evaluation research concerning the open practices badges, which may point to several issues in line with the quote above. Such as this research and paper drawing sub-optimal conclusions, and “hyping”:

        https://absolutelymaybe.plos.org/2017/08/29/bias-in-open-science-advocacy-the-case-of-article-badges-for-data-sharing/

        And there was the more recent paper of several of these “reformers” that seemed to have problems properly performing pre-registration if I remember correctly. This particular case has also been mentioned on this blog here:

        https://statmodeling.stat.columbia.edu/2024/09/26/whats-the-story-behind-that-paper-by-the-center-for-open-science-team-that-just-got-retracted/

        I think these two examples are relevant and possibly interesting examples concerning the quote above about meta-science. I would also like to note that I think it would make a possibly interesting and very useful topic concerning a possible paper, should Mr. Morey and/or others not have pondered this option.

        • Quote from the blog post: “Richard Morey wrote, “I think there’s an argument to be made that much meta-scientific work is a kind of mirror image of the empirical work it critiques,” and he shared this chart:”

          I thought of some more possible examples in line with this quote and what I already mentioned. Please check and verify, I am quickly writing things down and trying to remember things and go back to some things I wrote about earlier.

          Depending on what you view as “meta-scientific” and a “mirror image”, I thought of the following further examples. I think they more or less relate to, or involve, meta-science, and I think they can be seen as a mirror image or at least some vague reflection in some way, shape, or form in light of the quote.

          The above mentioned open practices badges that were the topic of some evaluation paper by some “reformers” that may have included some “hype” and sub-optimal conclusions are itself a mirror image, or a vague reflection, of an earlier effort by the journal Biostatistics (see Rowhani-Farid & Barnett, 2018, p. 3). The same goes for the Registered Reports, which can be seen as being a mirror image, or vague reflection, of efforts by The Lancet in 1997 (see Hardwicke & Iaonnidis, 2018).

          I wonder if the “redefine statistical significance” paper and it’s proposal are, in some way, a vague or clear reflection of earlier reactions to criticism of NHST by effectively largely ignoring the criticism. The proposal itself may also actually aggravate several biases cauesd by significance testing (see Amrhein & Greenland, 2018) and might make the replication crisis worse (see Crabe, 2017).

          If I understand things correctly, and they are still the case at this moment, Registered Reports explicitly have a role for journal editors and/or peer-reviewers to be involved in earlier phases of the research. In some way, I wonder if this is a mirror image of handing way too much power and influence to the journal-editor-peer-reviewer system which (arguably) have been (partly) responsible for many of the problematic issues.

          The registered reports investigation by Hardwicke & Ioannidis (2018) showed lack of transparency and/or problematic issues with pre-registration if I remember correctly, which has recently been replicated by some of these “reformers” if I understood correctly (see the blog post linked to above about the retracted paper). Seems to me to be at least a “vague reflection”, or perhaps even a “mirror image”, concerning these kinds of transparency and clarity issues regarding pre-registration.

          All this pre-print stuff was one of the things that got me enthusiastic around 2012 when certain proposals by some of these “reformers” were published. However, as time went by, I did not see that the “traditional” journal-editor-peer-review system went away or was being emphasized less. Quite the contrary, it seems to me that some these “reformers” are actively explicitly incorporating this “traditional” system into “new” proposals (e.g. see Registered Reports). Perhaps another example of something related to “meta-science” being a mirror image of (brushing aside or ignoring) problematic issues.

          I think some of the criticism regarding the now retracted paper by some of these “reformers” made clear that certain people seemed to “stand-by-their-open-science-man/woman”, and that less than optimal cliques or even cult-like behavior may have been displayed. That’s one of the reasons I have been disappointed with many of these “reformers” previous to this more recent situatoin. This more recent situation regarding the retracted paper again made clear to me that certain “group” or “insider” processes that may have facilitated problematic issues in the past may now also be present with some of these “reformers” and the “meta-science” they perform and promote and write about. All in all, I think that’s another mirror image, or at least a vague reflection.

          Talking about “reflection”, I have done enough of that for now. I just wanted to note a few more things in light of the quote by Mr. Morey, and in light of my comment that it might be an interesting and useful topic for a possible paper to consider.

  2. >The problem with bad science is not so much the overconfident conjectures–such steps may be psychologically necessary–so much as the unwillingness to reflect on contrary evidence, the unwillingness to admit error, and the practice of not confronting past mistakes.

    I think about the psychologically necessary part a lot – I do think it’s inherent to language that to convey information you have to sacrifice the uncertainty to some extent. It’s hard to find influential scientific ideas where the authors didn’t come out swinging. I agree with you that this is ok so long as we’re willing to update our views where they don’t pan out. But on a personal level, I find it uncomfortable, that if I speak about topics I’m an expert on with what feels like the appropriate level of caution, I risk just being ignored.

    • Quote from above: “But on a personal level, I find it uncomfortable, that if I speak about topics I’m an expert on with what feels like the appropriate level of caution, I risk just being ignored.”

      I am not an expert but try and be very careful in how I phrase things, and I notice this (or a lack of this) in other things I read by other writers. So, it may also be the case that, at least for some, phrasing things carefully and cautiously might have a different effect than being ignored. I view stating things cautiously as a possible sign of being a scientist with an (in my view) appropriate and important attitude or characteristic (or whatever term is most appropriate here).

      I also think that it might sometimes be the case that phrasing things in question form, like I did in the manuscript I quoted something from above, might actually lead to more (implicit) engagement or thought in the reader. Perhaps this all might go nicely with the fact that I wrote a manuscript with the title “Things I have wondered (so far)” in which I mostly just wonder about certain things.

      In that manuscript I also use the words “might” 52 times, “wonder” 49 times, and “could” 7 times (if I did the count-the-word analysis correctly just now), which could be a nice example of trying to be careful in stating things. And, there is always the option to not really care if one might possibly be ignored by some, or at a certain point in time, etc. Perhaps one can only do what one thinks might be best or wants to do. If I remember correctly, you even stated something like that in a previous comment on this blog when something similar was being discussed 5 or 6 months ago (?).

    • Have you read much of P.K. Feyerabend? He was fairly vocal about having to be a “big mouth” if you have anything deviant to say especially if you were not an established figure in the field.

      I say this mostly to suggest that it seems to have always been this way. In science, just like in most of life, it helps to sell. And unfortunately asserting confidence, even the unwarranted variety, is usually quite helpful in such endeavors. I hate it myself, but people being people, it’s probably here to stay. Best one can do is play the game while keeping in sharp focus that it is indeed a game. Once you start to drink your own Kool-Aid – as most do after a period of time – that’s when you’re really cooked!

    • On the path to truth, wisdom, and insight
      Sometimes all that may be needed is a shimmering light
      Something to illuminate, although it may not be very bright
      Just a little spark to help discern what might be wrong from what might be right

      Sometimes a faint whisper may make someone hear clearly for the first time
      Which may help them on their particular, individual climb
      Or it might make them wonder, or even stop on a dime
      And possibly turn around, and look for a path that is more sublime

      So, perhaps it’s hard to say
      What is “good” or “bad” in a certain way
      Glaring lights or loud sounds may lead one astray
      While a dim light or a whisper might sometimes lead the way

  3. The metascience and sociology of science stuff is interesting at first but gets old fast. The problem is enumerating badness, which only works when there are a few kinds of problematic practices to consider. In reality, there are effectively infinite ways to mess up a study (and get rewarded by significant differences for doing so):

    Why is “Enumerating Badness” a dumb idea? It’s a dumb idea because sometime around 1992 the amount of Badness in the Internet began to vastly outweigh the amount of Goodness. For every harmless, legitimate, application, there are dozens or hundreds of pieces of malware, worm tests, exploits, or viral code. Examine a typical antivirus package and you’ll see it knows about 75,000+ viruses that might infect your machine. Compare that to the legitimate 30 or so apps that I’ve installed on my machine, and you can see it’s rather dumb to try to track 75,000 pieces of Badness when even a simpleton could track 30 pieces of Goodness. In fact, if I were to simply track the 30 pieces of Goodness on my machine, and allow nothing else to run, I would have simultaneously solved the following problems:

    Spyware
    Viruses
    Remote Control Trojans
    Exploits that involve executing pre-installed code that you don’t use regularly
    Thanks to all the marketing hype around disclosing and announcing vulnerabilities, there are (according to some industry analysts) between 200 and 700 new pieces of Badness hitting the Internet every month. Not only is “Enumerating Badness” a dumb idea, it’s gotten dumber during the few minutes of your time you’ve bequeathed me by reading this article.

    The real solution is a white list of *good* practices. Ie, independent replication and testing otherwise surprising predictions.

    • Although, as a counter to that viewpoint, one might argue that various meta-science approaches do not seek to enumerate badness at the level of the “problematic practice”, but rather look to quantify emergent statistical phenomena that arise from whatever individual-level iniquities may ultimately contribute to some systematic divergence from what would be expected under a system of hypothetical perfect science (which would presumably be that emerging if everyone followed a white list of best practice).

  4. Re:the last row of the chart, I absolutely hate it when people say I have to provide an alternative to prove them wrong, otherwise they keep believing whatever BS they’re spouting. Even after I give them good counter-arguments. Your argument has been disproved, dude, get over it! This type of nonsense can be used as a rhetorical technique in any debate, not just in science.

  5. systematic divergence from what would be expected under a system of hypothetical perfect science

    Seems to be mostly divergence from what would be expected if the default null hypothesis (of no difference between groups) is true, which it is not.

Leave a Reply

Your email address will not be published. Required fields are marked *