Of course its preregistered. Just give me a sec

This is Jessica. I was going to post something on Bak-Coleman and Devezer’s response to the Protzko et al. paper on the replicability of research that uses rigor-enhancing practices like large samples, preregistration, confirmatory tests, and methodological transparency, but Andrew beat me to it. But since his post didn’t get into one of the surprising aspects of their analysis (beyond the paper making causal claim without a study design capable of assessing causality), I’ll blog on it anyway.

Bak-Coleman and Devezer describe three ways in which the measure of replicability that Protzko et al. use to argue that the 16 effects they study are more replicable than effects in prior studies deviates from prior definitions of replicability:

  1. Protzko et al. define replicability as the chance that any replication achieves significance in the hypothesized direction as opposed to whether the results of the confirmation study and the replication were consistent 
  2. They include self-replications in calculating the rate
  3. They include repeated replications of the same effect and replications across different effects in calculating the rate

Could these deviations in how replicability is defined have been decided post-hoc, so that the authors could present positive evidence for their hypothesis that rigor-enhancing practices work? If they preregistered their definition of replicability, we would not be so concerned about this possibility.  Luckily, the authors report that “All confirmatory tests, replications and analyses were preregistered both in the individual studies (Supplementary Information section 3 and Supplementary Table 2) and for this meta-project (https://osf.io/6t9vm).”

But wait – according to Bak-Coleman and Devezer:

the analysis on which the titular claim depends was not preregistered. There is no mention of examining the relationship between replicability and rigor-improving methods, nor even how replicability would be operationalized despite extensive descriptions of the calculations of other quantities. With nothing indicating this comparison or metric it rests on were planned a priori, it is hard to distinguish the core claim in this paper from selective reporting and hypothesizing after the results are known. 

Uh-oh, that’s not good. At this point, some OSF sleuthing was needed. I poked around the link above, and the associated project containing analysis code. There are a couple analysis plans: Proposed Overarching Analyses for Decline Effect final.docx, from 2018, and Decline Effect Exploratory analyses and secondary data projects P4.docx, from 2019. However, these do not appear to describe the primary analysis of replicability in the paper (the first describes an analysis that ends up in the Appendix, and the second a bunch of exploratory analyses that don’t appear in the paper). About a year later, the analysis notebooks with the results they present in the main body of the paper were added. 

According to Bak-Coleman on X/Twitter: 

We emailed the authors a week ago. They’ve been responsive but as of now, they can’t say one way or another if the analyses correspond to a preregistration. They think they may be in some documentation.

In the best case scenario where the missing preregistration is soon found, this example suggests that there are still many readers and reviewers for whom some signal of rigor suffices even when the evidence of it is lacking. In this case, maybe the reputation of authors like Nosek reduced the perceived need on the part of the reviewers to track down the actual preregistration. But of course, even those who invented rigor-enhancing practices can still make mistakes!

In the alternative scenario where the preregistration is not found soon, what is the correct course of action? Surely at least a correction is in order? Otherwise we might all feel compelled to try our luck at signaling preregistration without having to inconvenience ourselves by actually doing.

More optimistically, perhaps there are exciting new research directions that could come out of this. Like, wearable preregistration, since we know from centuries of research and practice that it’s harder to lose something when it’s sewn to your person. Or, we could submit our preregistrations to OpenAI, I mean Microsoft, who could make a ChatGPT-enabled Preregistration Buddy who not only trained on your preregistration, but also knows how to please a human judge who wants to ask questions about what it said.

18 thoughts on “Of course its preregistered. Just give me a sec

      • Jessica:

        I’d set epsilon to 2, just to be on the safe side. That would cover the notorious study where the estimated effect went in the opposite direction of the preregistration, but it was statistically significant at the 0.10 [sic] level, was published in a top journal, and received two awards. When the result is in the opposite direction of the preregistered claim, that’s epsilon=2, right?

  1. Quote from above:” In this case, maybe the reputation of authors like Nosek reduced the perceived need on the part of the reviewers to track down the actual preregistration. ”

    Maybe it was the intention of the group of collaborators to not only replicate “new discoveries” but also “old findings”. The possible missing preregistration, or other problems with it, might be a replication of previous behavior and actions and findings. From the top of my head, so please check and verify, I can clearly remember a paper by Hardwicke and Ioannidis (2018) that found similar issues with the preregistration of some “Registered Reports”, which if I’m not mistaken are heavily promoted by Nosek and many of his collaborators.

    I sometimes wonder how it is possible to mess this all up so badly, while almost everything you have done in the years prior to that is uttering sentences in which every 10th word is something like “transparency, “openness”, or “preregistration” or listening to others using these words. I just don’t understand it. But on the other hand that is consistent with my conclusion after following developments in Psychological Science in the past decade or so: I just don’t understand it.

    I can only understand it if it’s a way to somehow ultimately introduce some third-party control system somwhere down to road: “Look how incredibly difficult it is to write something down, and then do what it says, and check once in a while whether you are doing things according to what you’ve written down. We must have some institute or center verify and control this all. It’s all too much and too complicated and too difficult for the individual scientist!?”.

    I recently mentioned in a comment that I like wordplay in scientific writing and that I sometimes have a title for a paper in mind that involves this. One being “Psych101 or Psych1984?” that could be about things I find weird and/or absurd in Psychological Science. I also mentioned in that comment that I did not want to start writing because I would have to read certain things by certain people about certain topics that I don’t want to. The paper in question in this and yesterda’s blog would perhaps be a great candidate for the possible paper, but after havin browsed through it yesterday when I made a comment on yesterday’s blog I was reminded again of why I don’t want to read about certain things in certain papers by certain people anymore…

    • With “yesterdays blog” I mean the previous blogpost.

      This once in a while two-blogposts-a-day stuff messes with me keeping track of which blogpost was posted on which day I think…

    • I haven’t read the paper in question in detail (for reasons see above), but when I browsed through it I think I read something about pilot studies and exploratory research which is then subsequently replicated which all reminded me of an idea I posted on this exact blog back in 2017.

      I also remember mentioning the gist of the idea, if I am not mistaken, on some google discussion board before that, where if I remember correctly Mr. Nosek also commented and indicated that he was working on or thinking about something similar if I am not mistaken.

      I also posted the idea in several other places over the years.

      Anyway, I am linking to that idea here, should it be interesting for certain researchers that may want to think some more about, and perhaps try and experiment with, the basic idea of it all. The idea posted in 2017 differs in a few ways I think from the paper in question, which might be interesting for some people to read. I reason people can possibly compare and perhaps pick what they deem most useful:

      https://statmodeling.stat.columbia.edu/2017/12/17/stranger-than-fiction/#comment-628652

      • Quote from above: “I also remember mentioning the gist of the idea, if I am not mistaken, on some google discussion board before that, where if I remember correctly Mr. Nosek also commented and indicated that he was working on or thinking about something similar if I am not mistaken. ”

        I found the thread on the google discussion board. It is from 2013. Oh, how time flies. I posted as “fastchanges77” and started the thread with the basic outline of the idea I wrote about in more detail on this blog in 2017 linked to above.

        https://groups.google.com/g/openscienceframework/c/2nhHMdGGhrw

  2. So, what problem do people have with comparing intervals to define “replicated”? If they overlap, success.

    Whatever interval method was used for the original claim should also be used for the later ones. If you choose some methods yielding a wide interval, ok.

    That is a different problem to address than replication. Ie, it is easier to come up with alternative explanations.

    • Either way, there is still a popular scheme to get around these definitions of replication.

      Say study one shows that some intervention increases longevity, ie years of life by 20%.

      The replication study also measures longevity, but now that means the height of the person, or length of the mouses tail, etc.

      That is an extreme example but its what indirect replications amount to.

  3. It would have been good to report the replicability measure also with another more restrictive operational definition (for example, without counting the self-replications and considering in a different way the multiple replications of the same effect). Hopefully, some of this rigor enhancing practices won’t become a “Green-washing” type of thing. A bit like some way too optimistic a priori power analyses sometimes done because everyone does them and they are well seen by journals rather than reflecting a serious reasoning on one’s own study design.

Leave a Reply

Your email address will not be published. Required fields are marked *