In the context of our recent discussion of the p-curve paper, Richard Morey wrote, “I think there’s an argument to be made that much meta-scientific work is a kind of mirror image of the empirical work it critiques,” and he shared this chart:

I think Morey is on to something here, but, as someone who does a lot of empirical science and a lot of meta-science, I think there’s one big thing he’s missing, one major asymmetry between empirical science and meta-science, and that is that bad empirical science makes strong claims, and the role of meta-science is to question the evidential support behind these claims, not usually to make a positive claim in itself.
The usual pattern goes like this: empirical scientists collect data D, perform analysis A, and use these to make strong general claim X about the world. The meta-scientist then comes along to assess the evidence. A negative meta-science analysis comes to the conclusion that D + A do not provide good evidence for X. The meta-science analysis does not make the strong claim that X is false, let alone the even stronger claim that some preferred alternative Y is true.
This comes up all the time. Some Cornell psychology professor claims to have strong evidence for extra-sensory perception or influence of food labeling on eating or whatever. The meta-scientist comes along and notes irregularities with the data or analysis and provides an alternative story of how these apparently convincing patterns in data could have come to be. The conclusion of the meta-scientific report is not that ESP or large effects of food labels don’t exist but rather that the published record does not provide good evidence of these extraordinary claims. (And indeed the claims are extraordinary, which is how they got so much publicity in the first place.)
It’s the all-important distinction between truth and evidence. I know that Morey understands this distinction and I’m not saying that anything in his above chart is wrong; I’m just trying to put it in the larger perspective of scientific inquiry.
In discussing the above asymmetry between empirical science and meta-science, I’m not saying that meta-science is better. Meta-science is fundamentally parasitic on empirical science, and, sure, empirical science is associated with bold claims, but it’s through making bold leaps–and being willing to retract those leaps as needed–that we make progress. The problem with bad science is not so much the overconfident conjectures–such steps may be psychologically necessary–so much as the unwillingness to reflect on contrary evidence, the unwillingness to admit error, and the practice of not confronting past mistakes.
And also the really stupid things that people say and never apologize for.
For the casual reader, this post is referencing a discussion of this paper
https://www.tandfonline.com/doi/full/10.1080/01621459.2025.2544397
and this response
https://bsky.app/profile/did:plc:c56dfqq2kropq4ys65h5pvgr/post/3lznpffulh22s
that took place in September on this blog
https://statmodeling.stat.columbia.edu/2025/09/25/on-the-poor-statistical-properties-of-the-p-curve-meta-analytic-procedure/
Link added; thanks.
Quote from the blog post: “I think there’s an argument to be made that much meta-scientific work is a kind of mirror image of the empirical work it critiques,”
Ah, that’s reassuring to me to read in the sense that I may not be the only one to think about possible problematic issues concerning “meta-science”. At a certain point in time I had come across multiple things that seem to me to possibly be “conceptual replications” of problematic issues. Some of these were proposed as being “solutions” or “improvements” in light of the so-called “replication crisis”. I thought it might be useful and/or amusing to make a manuscript about all of these things, in which I also ask many questions. Some of these questions might resonate with the topic of this blog post:
“Is there a risk that “meta-science” and/or evaluation research can be used (in the short-term) to “nudge” or “steer” (parts of) Psychological Science in a certain direction without much and/or sufficient thought concerning the soundness, desirability, and/or validity (in the long-term) of the object or process under investigation? Is there a risk that the prospect of future “meta-scientific” and/or evaluation research can be used to ignore, brush aside, and/or neglect (possibly and/or probably) valid criticism? Can “meta-scientific” and/or evaluation research have similar problems as “normal” research concerning its validity, conflicts of interest, possible corruption, soundness of design and conclusions, etc.?”
Anon:
Yes, one thing I’ve discussed in the past is the way that the meta-science movement in psychology has been an awkward coalition of reformers who wanted to use meta-science to discredit a bunch of old research, and status-quo types who wanted to use meta-science as a way to bolster the credibility of that old work. Both groups were reacting to the replication crisis but in different ways, and I think it’s been difficult for Brian Nosek and others to keep the support of both factions.
Some of these “reformers” might provide great examples concerning the quote “I think there’s an argument to be made that much meta-scientific work is a kind of mirror image of the empirical work it critiques,”
If I am not mistaken, there was a blog post by Bastian in 2017 for instance about evaluation research concerning the open practices badges, which may point to several issues in line with the quote above. Such as this research and paper drawing sub-optimal conclusions, and “hyping”:
https://absolutelymaybe.plos.org/2017/08/29/bias-in-open-science-advocacy-the-case-of-article-badges-for-data-sharing/
And there was the more recent paper of several of these “reformers” that seemed to have problems properly performing pre-registration if I remember correctly. This particular case has also been mentioned on this blog here:
https://statmodeling.stat.columbia.edu/2024/09/26/whats-the-story-behind-that-paper-by-the-center-for-open-science-team-that-just-got-retracted/
I think these two examples are relevant and possibly interesting examples concerning the quote above about meta-science. I would also like to note that I think it would make a possibly interesting and very useful topic concerning a possible paper, should Mr. Morey and/or others not have pondered this option.
Quote from the blog post: “Richard Morey wrote, “I think there’s an argument to be made that much meta-scientific work is a kind of mirror image of the empirical work it critiques,” and he shared this chart:”
I thought of some more possible examples in line with this quote and what I already mentioned. Please check and verify, I am quickly writing things down and trying to remember things and go back to some things I wrote about earlier.
Depending on what you view as “meta-scientific” and a “mirror image”, I thought of the following further examples. I think they more or less relate to, or involve, meta-science, and I think they can be seen as a mirror image or at least some vague reflection in some way, shape, or form in light of the quote.
The above mentioned open practices badges that were the topic of some evaluation paper by some “reformers” that may have included some “hype” and sub-optimal conclusions are itself a mirror image, or a vague reflection, of an earlier effort by the journal Biostatistics (see Rowhani-Farid & Barnett, 2018, p. 3). The same goes for the Registered Reports, which can be seen as being a mirror image, or vague reflection, of efforts by The Lancet in 1997 (see Hardwicke & Iaonnidis, 2018).
I wonder if the “redefine statistical significance” paper and it’s proposal are, in some way, a vague or clear reflection of earlier reactions to criticism of NHST by effectively largely ignoring the criticism. The proposal itself may also actually aggravate several biases cauesd by significance testing (see Amrhein & Greenland, 2018) and might make the replication crisis worse (see Crabe, 2017).
If I understand things correctly, and they are still the case at this moment, Registered Reports explicitly have a role for journal editors and/or peer-reviewers to be involved in earlier phases of the research. In some way, I wonder if this is a mirror image of handing way too much power and influence to the journal-editor-peer-reviewer system which (arguably) have been (partly) responsible for many of the problematic issues.
The registered reports investigation by Hardwicke & Ioannidis (2018) showed lack of transparency and/or problematic issues with pre-registration if I remember correctly, which has recently been replicated by some of these “reformers” if I understood correctly (see the blog post linked to above about the retracted paper). Seems to me to be at least a “vague reflection”, or perhaps even a “mirror image”, concerning these kinds of transparency and clarity issues regarding pre-registration.
All this pre-print stuff was one of the things that got me enthusiastic around 2012 when certain proposals by some of these “reformers” were published. However, as time went by, I did not see that the “traditional” journal-editor-peer-review system went away or was being emphasized less. Quite the contrary, it seems to me that some these “reformers” are actively explicitly incorporating this “traditional” system into “new” proposals (e.g. see Registered Reports). Perhaps another example of something related to “meta-science” being a mirror image of (brushing aside or ignoring) problematic issues.
I think some of the criticism regarding the now retracted paper by some of these “reformers” made clear that certain people seemed to “stand-by-their-open-science-man/woman”, and that less than optimal cliques or even cult-like behavior may have been displayed. That’s one of the reasons I have been disappointed with many of these “reformers” previous to this more recent situatoin. This more recent situation regarding the retracted paper again made clear to me that certain “group” or “insider” processes that may have facilitated problematic issues in the past may now also be present with some of these “reformers” and the “meta-science” they perform and promote and write about. All in all, I think that’s another mirror image, or at least a vague reflection.
Talking about “reflection”, I have done enough of that for now. I just wanted to note a few more things in light of the quote by Mr. Morey, and in light of my comment that it might be an interesting and useful topic for a possible paper to consider.
>The problem with bad science is not so much the overconfident conjectures–such steps may be psychologically necessary–so much as the unwillingness to reflect on contrary evidence, the unwillingness to admit error, and the practice of not confronting past mistakes.
I think about the psychologically necessary part a lot – I do think it’s inherent to language that to convey information you have to sacrifice the uncertainty to some extent. It’s hard to find influential scientific ideas where the authors didn’t come out swinging. I agree with you that this is ok so long as we’re willing to update our views where they don’t pan out. But on a personal level, I find it uncomfortable, that if I speak about topics I’m an expert on with what feels like the appropriate level of caution, I risk just being ignored.
Quote from above: “But on a personal level, I find it uncomfortable, that if I speak about topics I’m an expert on with what feels like the appropriate level of caution, I risk just being ignored.”
I am not an expert but try and be very careful in how I phrase things, and I notice this (or a lack of this) in other things I read by other writers. So, it may also be the case that, at least for some, phrasing things carefully and cautiously might have a different effect than being ignored. I view stating things cautiously as a possible sign of being a scientist with an (in my view) appropriate and important attitude or characteristic (or whatever term is most appropriate here).
I also think that it might sometimes be the case that phrasing things in question form, like I did in the manuscript I quoted something from above, might actually lead to more (implicit) engagement or thought in the reader. Perhaps this all might go nicely with the fact that I wrote a manuscript with the title “Things I have wondered (so far)” in which I mostly just wonder about certain things.
In that manuscript I also use the words “might” 52 times, “wonder” 49 times, and “could” 7 times (if I did the count-the-word analysis correctly just now), which could be a nice example of trying to be careful in stating things. And, there is always the option to not really care if one might possibly be ignored by some, or at a certain point in time, etc. Perhaps one can only do what one thinks might be best or wants to do. If I remember correctly, you even stated something like that in a previous comment on this blog when something similar was being discussed 5 or 6 months ago (?).
Have you read much of P.K. Feyerabend? He was fairly vocal about having to be a “big mouth” if you have anything deviant to say especially if you were not an established figure in the field.
I say this mostly to suggest that it seems to have always been this way. In science, just like in most of life, it helps to sell. And unfortunately asserting confidence, even the unwarranted variety, is usually quite helpful in such endeavors. I hate it myself, but people being people, it’s probably here to stay. Best one can do is play the game while keeping in sharp focus that it is indeed a game. Once you start to drink your own Kool-Aid – as most do after a period of time – that’s when you’re really cooked!
On the path to truth, wisdom, and insight
Sometimes all that may be needed is a shimmering light
Something to illuminate, although it may not be very bright
Just a little spark to help discern what might be wrong from what might be right
Sometimes a faint whisper may make someone hear clearly for the first time
Which may help them on their particular, individual climb
Or it might make them wonder, or even stop on a dime
And possibly turn around, and look for a path that is more sublime
So, perhaps it’s hard to say
What is “good” or “bad” in a certain way
Glaring lights or loud sounds may lead one astray
While a dim light or a whisper might sometimes lead the way
The metascience and sociology of science stuff is interesting at first but gets old fast. The problem is enumerating badness, which only works when there are a few kinds of problematic practices to consider. In reality, there are effectively infinite ways to mess up a study (and get rewarded by significant differences for doing so):
The real solution is a white list of *good* practices. Ie, independent replication and testing otherwise surprising predictions.
Although, as a counter to that viewpoint, one might argue that various meta-science approaches do not seek to enumerate badness at the level of the “problematic practice”, but rather look to quantify emergent statistical phenomena that arise from whatever individual-level iniquities may ultimately contribute to some systematic divergence from what would be expected under a system of hypothetical perfect science (which would presumably be that emerging if everyone followed a white list of best practice).
The Anna Karenina principle of bad science:
“Every good piece of research is alike. Every bad piece is bad in its own way.”
Re:the last row of the chart, I absolutely hate it when people say I have to provide an alternative to prove them wrong, otherwise they keep believing whatever BS they’re spouting. Even after I give them good counter-arguments. Your argument has been disproved, dude, get over it! This type of nonsense can be used as a rhetorical technique in any debate, not just in science.
Seems to be mostly divergence from what would be expected if the default null hypothesis (of no difference between groups) is true, which it is not.