I was having a discussion with someone about problems with the science reform movement (as discussed here by Jessica), and he shared his opinion that “Scientific reform in some corners has elements of millenarian cults. In their view, science is not making progress because of individual failings (bias, fraud, qrps) and that if we follow a set of rituals (power analysis, preregistration) devised by the leaders than we can usher in a new era where the truth is revealed (high replicability).”
My quick reaction was that this reminded me of an annoying thing where people use “religion” as a term of insult. When this came up before, I wrote that maybe it’s time to retire use of the term “religion” to mean “uncritical belief in something I disagree with.”
But then I was thinking about this all from another direction, and I think there’s something there there. Not the “millenarian cults” thing, which I think was an overreaction on my correspondent’s part.
Rather, I see a paradox. From his perspective, my correspondent sees the science reform movement as having a narrow perspective, an enforced conformity that leads it into unforced errors such as publishing a high-profile paper promoting preregistration without actually itself following preregistered analysis plans. OK, he doesn’t see all of the science reform movement as being so narrow—for one thing, I’m part of the science reform movement and I wasn’t part of that project!—but he seems some core of the movement being stuck in narrow rituals and leader-worship.
But I think it’s kind of the opposite. From my perspective, the core of the science reform movement (the Open Science Framework, etc.) has had to make all sorts of compromises with conservative forces in the science establishment, especially within academic psychology, in order to keep them on board. To get funding, institutional support, buy-in from key players, . . . that takes a lot of political maneuvering.
I don’t say this lightly, and I’m not using “political” as a put-down. I’m a political scientist, but personally I’m not very good at politics. Politics takes hard work, requiring lots of patience and negotiation. I’m impatient and I hate negotiation; I’d much rather just put all my cards face-up on the table. For some activities, such as blogging and collaborative science, these traits are helpful. I can’t collaborate with everybody, but when the connection’s there, it can really work.
But there’s more to the world than this sort of small-group work. Building and maintaining larger institutions, that’s important too.
So here’s my point: Some core problems with the open-science movement are not a product of cult-like groupthink. Rather, it’s the opposite: this core has been structured out of a compromise with some groups within psychology who are tied to old-fashioned thinking, and this politically-necessary (perhaps) compromise has led to some incoherence, in particular the attitude or hope that, by just including some preregistration here and getting rid of some questionable research practices there, everyone could pretty much continue with business as usual.
Summary
The open-science movement has always had a tension between burn-it-all-down and here’s-one-quick-trick. Put them together and it kinda sounds like a cult that can’t see outward, but I see it as more the opposite, as an awkward coalition representing fundamentally incoherent views. But both sides of the coalition need each other: the reformers need the old institutional powers to make a real difference in practice, and the oldsters need the reformers because outsiders are losing confidence in the system.
The good news
The good news for me is that both groups within this coalition should be able to appreciate frank criticism from the outside (they can listen to me scream and get something out of it, even if they don’t agree with all my claims) and should also be able to appreciate research methods: once you accept the basic tenets of the science reform movement, there are clear benefits to better measurement, better design, and better analysis. In the old world of p-hacking, there was no real reason to do your studies well, as you could get statistical significance and publication with any old random numbers, along with a few framing tricks. In the new world of science reform—even imperfect science reform, this sort of noise mining isn’t so effective, and traditional statistical ideas of measurement, design, and analysis become relevant again.
So that’s one reason I’m cool with the science reform movement. I think it’s in the right direction: its dot product with the ideal direction is positive. But I’m not so good at politics so I can’t resist criticizing it too. It’s all good.
Reactions
I sent the above to my correspondent, who wrote:
I don’t think it is a literal cult in the sense that carries the normative judgments and pejorative connotations we usually ascribe to cults and religions. The analogy was more of a shorthand to highlight a common dynamic that emerges when you have a shared sense of crisis, ritualistic/procedural solutions, and a hope that merely performing these activities will get past the crisis and bring about a brighter future. This is a spot where group-think can, and at times possibly should, kick in. People don’t have time to each individually and critically evaluate the solutions, and often the claim is that they need to be implemented broadly to work. Sometimes these dynamics reflect a real problem with real solutions, sometimes they’re totally off the rails. All this is not to say I’m opposed to scientific reform; I’m very much for it in the general sense. There’s no shortage of room for improvement in how we turn observations into understanding, from improving statistical literacy and theory development to transparency and fostering healthier incentives. I am, however, wary of the uncritical belief that the crisis is simply one of failed replications and that the performance of “open science rituals” is sufficient for reform, across the breadth of things we consider science. As a minor point, I don’t think many of the vast majority of prominent figures in open science intend for these dynamics to occur, but I do think they all should be wary of them.
There does seem to be a problem that many researchers are too committed to the “estimate the effect” paradigm and don’t fully grapple with the consequences of high variability. This is particularly disturbing in psychology, given that just about all psychology experiments study interactions, not main effects. Thus, a claim that effect sizes don’t vary much is a claim that effect sizes vary a lot in the dimension being studied, but have very little variation in other dimensions. Which doesn’t make a lot of sense to me.
Getting back to the open-science movement, I want to emphasize the level of effort it takes to conduct and coordinate these big group efforts, along with the effort required to keep together that the coalition of skeptics (who see preregistration as a tool for shooting down false claims) and true believers (who see preregistration as a way to defuse skepticism about their claims) and get these papers published in top journals. I’d also say it takes a lot of effort for them to get funding, but that would be kind of a cheap shot, given that I too put in a lot of effort to get funding!
Anyway, to continue, I think that some of the problems with the science reform movement are that it effectively promises different things to different people. And another problem is with these massive projects that inevitably include things that not all the authors will agree with.
So, yeah, I have a problem with simplistic science reform prescriptions, for example recommendations to increase sample size without any nod toward effect size and measurement. But much much worse, in my opinion, are the claims of success we’ve seen from researchers and advocates who are outside the science-reform movement. I’m thinking here about ridiculous statements such as the unfounded claim of 17 replications of power pose, or the endless stream of hype from the nudgelords, or the “sleep is your superpower” guy, or my personal favorite, the unfounded claim from Harvard that “the replication rate in psychology is quite high—indeed, it is statistically indistinguishable from 100%.”
It’s almost enough to stop here with the remark that the scientific reform movement has been lucky in its enemies.
But I also want to say that I appreciate that the “left wing” of the science reform movement—the researchers who envision replication and preregistration and the threat of replication and preregistration as a tool to shoot down bad studies—have indeed faced real resistance within academia and the news media to their efforts, as lots of people will hate the bearers of bad news. And I also appreciate that the “right wing” of the science reform movement—the researchers who envision replication and preregistration as a way to validate their studies and refute the critics—in that they’re willing to put their ideas to the test. Not always perfectly, but you have to start somewhere.
While I remain annoyed at certain aspects of the mainstream science reform movement, especially when it manifests itself in mass-authored articles such as the notorious recent non-preregistered paper on the effects of preregistration, or that “Redefine statistical significance” article, or various p-value hardliners we’ve encountered over the decades, I also respect the political challenges of coalition-building that are evident in that movement.
So my plan remains to appreciate the movement while continuing to criticize its statements that seem wrong or do not make sense.
I sent the above to Jessica Hullman, who wrote:
I can relate to being surprised by the reactions of open science enthusiasts to certain lines of questioning. In my view, how to fix science is as about a complicated question as we will encounter. The certainty/level of comfortableness with making bold claims that many advocates of open science seem to have is hard for me to understand. Maybe that is just the way the world works, or at least the way it works if you want to get your ideas published in venues like PNAS or Nature. But the sensitivity to what gets said in public venues against certain open science practices or people reminds me very much of established academics trying to hush talk about problems in psychology, as though questioning certain things is off limits. I’ve been surprised on the blog for example when I think aloud about something like preregistration being imperfect and some commenters seem to have a visceral negative reaction to see something like that written. To me that’s the opposite of how we should be thinking.
As an aside, someone I’m collaborating with recently described to me his understanding of the strategy for getting published in PNAS. It was 1. Say something timely/interesting, 2. Don’t be wrong. He explained that ‘Don’t be wrong’ could be accomplished by preregistering and large sample size. Naturally I was surprised to hear #2 described as if it’s really that easy. Silly me for spending all this time thinking so hard about other aspects of methods!
The idea of necessary politics is interesting; not what I would have thought of but probably some truth to it. For me many of the challenges of trying to reform science boil down to people being heuristic-needing agents. We accept that many problems arise from ritualistic behavior, but we have trouble overcoming that, perhaps because no matter how thoughtful/nuanced some may prefer to be, there’s always a larger group who want simple fixes / aren’t incentivized to go there. It’s hard to have broad appeal without being reductionist I guess.
I’m surprised your correspondent didn’t cite Feynman on “cargo cult science”. He even specified psychological studies as fitting that description.
Just being a matter of replications does leave out certain things. There isn’t really a concept of “replication” in history, but that is the idea Anton Howes at Age of Invention had to borrow when he asked “Does History Have a Replication Crisis?”*. And he’s only gotten more pessimistic since then, as the journal “History & Technology” as well as the British Society for the History of Science have both rallied around the paper numerous people have pointed out is incorrect, denouncing the critics for having bad motives in criticizing it. The former did at least acknowledge one factual error, but treated that as not mattering. That probably sounds familiar to you, but their rather open denial of a single standard for empiricism goes beyond what I typically see you complaining about in social science, and since there will never be a failed “replication” per se they might just continue insisting they were right.
* https://www.ageofinvention.xyz/p/age-of-invention-does-history-have
These definitions are fine, but if you read on the author starts conflating “same result” with same interpretation/conclusion.
In history, the “result” being reproduced would be stuff like a translation, or that some document can be found in a given archive, or a carbon date in a database, etc.
A replication would be finding another copy of the document from a similar or earlier era, redoing the carbon dating, and so on.
It is perfectly fine to then have competing interpretations of that evidence, which we then compare using bayes rule.
And distinguishing between various interpretations is the hard part. The orthagonal issue of replication/reproduction is supposed to be the easy part addressed by basic scholarly behavior.
There really is no excuse for there to be such huge problems with replication in psych and medical research. It is a purely cultural problem with offering and seeking funding for independent direct replications. The easiest fix is to fund half the novel studies and make replication (by someone else) a standard aspect of the grant. This would then show we can communicate the important factors required to generate similar data.
But, everytime that is tried the results are so bad I think funding and policy-setting agencies are just scared of what it means.
As someone who has followed, and even involved with, the reform/open science movement I want to note that I very much resonate with sentences like:
“I’ve been surprised on the blog for example when I think aloud about something like preregistration being imperfect and some commenters seem to have a visceral negative reaction to see something like that written.”
It is these exact kinds of things, in combination with several of the possibly “strange” papers “(…) such as the notorious recent non-preregistered paper on the effects of preregistration, or that “Redefine statistical significance” article,(…)”, and the in my opinion too much groupthink/cult-like-sfuff that made me wonder whether I should still be involved with all this, and even made me stop with it all.
I have lost faith.
I think many problematic issues might not really be solved, and are possibly “conceptually” replicated, and certain proposals might even make things worse. For those interested in some more details and information and reasons concerning this all, I would like to refer to a manuscript I wrote titled “Psychological Science Replicates Just Fine, Thanks” which can be found on SSRN.
Slightly off-topic, but I refer people to your article with Hal Stern on “The difference in stat sig and not is not itself significant” in peer reviews maybe one out of three articles I review.
It is just a crazy common mistake (at least in the social science papers I review). I just find it so much easier to reason about variability when you are specific about your estimator, so it just goes to `[Estimate1 – Estimate2] (SE_Dif)`, and then you can say “well, experiment 1 was quite noisy, so SE_Dif is large”. (I do find the quote about 100% quite lulz as well.)
Or SE_Dif is small, so even small differences (which are reasonable to expect) would be detected, but probably doesn’t falsify the general theory people are interested in.
In my view, one of the biggest challenges to science reform is that it actually was not the practices of science that were the problem. The problem was and is that many people don’t understand the meaning behind the practices of science in the first place—they were already a “cargo cult” practicing “rituals”. Andrew has discussed this on the blog in the context of measurement and study design. I would describe the problem as a lack of theory, where “theory” is a specification of how putative constructs lead to observable outcomes along with the mapping between measurements/manipulations and those constructs. Without that kind of theory, even just as a provisional theory, the results of studies are devoid of scientific meaning.
For example, I think that testing a null hypothesis by computing a p value can be valuable in some circumstances. If you have a substantive theory that says what the “null hypothesis” is and why it makes the predictions it does, then it is often useful to know how unusual a particular result is. If I have a theory that says outcome A should occur 80% of the time, and I get a result where A occurred 85% of the time, how much of a discrepancy is that? Can my theory still accommodate that result or is something else going on? Can something else explain my new result while still explaining the entire history of prior results to that point? These are all examples of meaningful questions that could be addressed with existing (and old) practices of science—the problem is that too many people weren’t asking those questions because they didn’t know that’s what those practices referred to.
For me, this is why many of the reforms that focus on the practices of science fall flat. I don’t see them as addressing the fundamental educational problem that we are training people to perform rituals without knowing why. So it feels like many reforms have the attitude, “science is too hard for people to actually do, so let’s just give them some checklists to follow instead.” Whether reformists actually hold that attitude or not, that’s how it often comes across. And this perception leads some people to feel insulted (“hey, I know what I’m doing!”) and others (including me) to view many reforms as nihilistic or cynical.
That said, I still have hope that reforming the training of scientists will ultimately improve the enterprise of science. Unfortunately, large-scale reforms of that type are not only practically difficult, they are (as Andrew points out) politically difficult as well.
For any real-world problem, it is possible for the critics to correctly identify what is currently wrong without their proposals of how to fix it being correct.
However, because of the way our world is structured, where careers are realistically about 30–35 years long and everyone has to eat, it will almost inevitably become necessary for reformers (particularly those who have been given grants or other incentives to advance the reform process) to claim that their efforts are making a difference after a relatively short period — say, 5 years, maybe 10, but not much more — even if the justification of that claim requires a great deal of motivated reasoning. This is standard in politics (“After 20 years of misgovernment by Party X, we in Party Y have made great strides in just 4 years, and we just need you to vote us back in one more time to finish the job”), but I think it is likely to end up applying to the science reform movement as well.
Of course, I am in the privileged position, as a latecomer in life to the science racket, that I have no skeletons in my cupboard from my time as a functionary of Party X, but also that I don’t need to hitch my putting-food-on-the-table wagon to the efforts of Party Y. The downside of that position is that I will not be around for too much longer. (This is not an announcement of any unfortunate life-shortening medical condition, but at 63 I am aware that my remaining days on this planet are unlike to exceed 10,000!)
Nick:
Interesting. Maybe my perspective is different than some others because I can realistically expect to have a 50-year career.
“…and the oldsters need the reformers because outsiders are losing confidence in the system”
I don’t think the reformers are needed to improve confidence. Only the perception of openness is necessary, which I think is why we see a lot of journals with superficial unenforced open data/materials policies. They don’t need research to be more reliable or valid. They just want the facade of reform.
I know this view is cynical, but it matches most of my experience. Far too many publications claim to have open data/materials, but then don’t. Far too many journals won’t retract or publish a commentary when there are glaring statistical errors. And far too many authors respond to concerns by attacking the people who point out those concerns.
“Far too many publications claim to have open data/materials, but then don’t.”
Are there any examples of sanctions for dishonestly claiming to have embraced reforms? Even something along the lines of a correction noting the data and code are not actually available.
Not that I’ve ever seen. I know of at least one paper that claimed the data was open and then issued a correction that it is available on request… and then the authors ignored my requests.
This post overlooks the significant changes in the political landscape affecting scientific reform over the past decade. Initially, advocates of scientific reform faced significant resistance in promoting basic standards of transparency. Now, a decade later, these practices have become mainstream, with early reformers now in influential positions such as professors, journal editors, and leaders in their fields. However, those aiming to critique or further reform these practices encounter similar political hurdles that early reformers faced in the 2010s.
Consider the response to Szollosi’s 2019 paper, both on Twitter and in subsequent literature. The scientific reform community’s reaction on Twitter was often harsh, featuring insults and mockery. Despite being cited over 100 times in the past three years, this paper has seen limited engagement from prominent reform advocates. Its citation network reveals a group of scientists focused on reforming the reform movement, yet their work receives minimal attention within the mainstream reform community (Jessica and Andrew are exceptions), indicating a one-sided citation dynamic. The tendency towards public mockery with limited serious academic engagement resembles what reformers faced a decade back (“methodological terrorism!”).
I agree with Andrew that the term ‘cult’ is not constructive and a political lens offers a more suitable approach. Adopting it, we still need to grapple with the essence of the ‘cult’ critique. Maintaining scientific reform’s political wins has resulted in rigidity and unwillingness to change that often resembles the previous system.
Anon:
I had not heard of Szollosi’s 2019 paper so I took a look just now. It’s “Arrested Theory Development: The Misguided Distinction Between Exploratory and Confirmatory Research,” by Aba Szollosi and Chris Donkin. Ironically (to me), it’s published in Perspectives on Psychological Science, a journal that a couple years previously had published an unfounded attack on me that their editorial board very rudely refused to correct. (And, unrelated to me, as late as 2021 the Association for Psychological Science was promoting pseudoscience such as an absolutely horrible “lucky golf ball” study from their archives. Very little capacity for embarrassment on their part.)
I agree with much of what Szollosi and Donkin say, in particular the bit about preregistration and other procedural steps being overrated. Indeed, just this month, science reformer Uri Simonsohn wrote, “Pre-registration is the best and possibly only solution to p-hacking.” I’m a big fan of Simonsohn, but I disagree with that statement of his, partly because there are other solutions to p-hacking (for example, there’s multilevel modeling) and also because I’m not a huge fan of the term “p-hacking” as it seems to me to imply intentionality, and I think a lot of bad statistical analyses, including unadjusted multiple potential comparisons (“forking paths”) arise are not done on purpose.
The point about intentionality is relevant here, in that sins such as “p-hacking,” “questionable research practices,” “harking,” etc., are all things that are easy to fix—just stop doing these bad things!—, and many of the proposed solutions, such as preregistration and increased sample size, require some effort but no thought.
As I’ve been saying for a decade now (!), preregistration has a valuable indirect function of making it more difficult to do bad science. It does not directly turn bad science into good science. That doesn’t make preregistration a bad idea—recently I’ve been preregistering studies and, more generally, simulating data before gathering any data (see here and here)—; we should just be aware that this sort of procedural step can only one small part of the story. Ultimately, science is about the substance of science, not just about the scientific method, and I think that’s a key message of the Szollosi and Donkin paper.
There’s something interesting here, though, that links the two perspectives. If you do things right, your preregistration will involve the substance of what you’re studying and will not merely be a procedural step, a form of paperwork that exists to validate the p-values that your study will produce. Rather, doing this preregistration will require simulating fake data, which in turn will require hypothesizing a full model of the underlying process.
I recognize that what I just described is not the usual thing that is meant by “preregistration,” which is more along the lines of: “We will perform this comparison and use a 2-sided test,” etc. But it could be! I think this is a useful connection.
Returning to the Szollosi and Donkin paper:
– I agree with what they say about “The Misguided Exploratory–Confirmatory Distinction”; see this post which actually makes this point in a discussion of a different Donkin paper.
– Their point that “experimental tests are often superfluous” reminds me of my argument that certain studies are “dead on arrival,” that we can know they are hopeless even before seeing the data. My favorite example here is that claim, based on a sample of 3000 people, that beautiful parents are more likely to have girl babies. Studying this with any realistic effect size would require something on the order of a million people. My only point here is that my argument was statistical, based on comparison of effect size and variation, whereas Szollosi and Donkin are reasoning on a more theoretical level.
– Their statement, “the aim of science is to create explanations that are inflexible,” reminds me of my conception of Bayesian data analysis, applied in our book and derived originally from my reading of E. T. Jaynes, of a statistical workflow that proceeds by making strong assumptions, then using data to find problems with these assumptions, then using this information to improve our models.
So, yeah, I like Szollosi and Donkin’s argument. Maybe it’s a good thing that I’m not on twitter so I didn’t see how it was being insulted and mocked. I did a quick twitter search of *Szollosi and Donkin* but only saw positive comments, so maybe it had a better reaction than you’re remembering!
That all said, it does seem to me that direct replication has a valuable function within the scientific community, as being a way to convince some people that, yeah, a certain published finding is artifactual (i.e., noise mining). Theoretical and statistical analysis should be enough to show that some studies are dead on arrival, just too noisy to detect any realistically-sized effect—but there’s nothing like a failed replication to make the point. I’m not saying that a failed replication will convince everybody—researchers will notoriously claim that non-replications are actually replications (examples include the ESP researcher claiming, as a replication of his study from 2011, a paper published several years before; the ovulation-and-clothing researchers doing their own failed replication and declaring it a success by adding a previously undiscussed claim of interaction with weather; and one of the power-pose researcher claiming “17 replications” even though there weren’t)—, but I think even the threat of failed replication can help things. I think that one reason some established researchers were so afraid of the replication movement is that, before everyone was talking about replications, it was seemingly impossible to ever knock a published claim off its perch of assumed correctness. Just the idea you could have a failed replication, that’s a step forward.
I didn’t know about the 2019 paper; I had the 2020 paper ‘Is Preregistration Worthwhile?’ in mind. It’s a bit embarrassing, even anon. Regardless, I appreciate your response. Very enlightening.
Ok, now checking, it seems that the Szollosi and Donkin paper is from 2021. It’s just the first link that popped up when I googled *Szollosi 2019 paper*.
Searching *Is Preregistration Worthwhile?* on Google scholar is hilarious, as it gives the following links:
Is preregistration worthwhile?
Preregistration is hard, and worthwhile
What should a preregistration contain?
Preregistration is redundant, at best
Does preregistration improve the credibility of research findings?
A survey on how preregistration affects the research workflow: Better science but more work
Preregistration: the good, the bad, and the confusing
Why preregistration makes me nervous
It just goes on and on. Very meta.
Perhaps this is obvious, but: “science” covers a very wide range of topics and that issues like pre-registration and forking paths are much more important in some areas than in others. The types of ‘science reform’ that are most needed in social science or political science are not necessarily the same as those needed in climatology or ecology or biology or cultural anthropology.
+1
Couldn’t agree more. It’s hard to see replication, preregistration and open data as viable reforms for science that requires decade(s) long studies in, say, field ecology. Yet you’d be surprised how uphill of a battle it is to publish this argument. It felt like one of the reviewers was going to hit caps lock at any moment and let it rip.
What’s to stop someone preregistering an analytical plan to investigate a particular ecological phenomenon in an existing long-term dataset? Or to stop that dataset being open? Government funded long-term ecological monitoring schemes in the UK typically have openness as part of their funding agreement.
Don’t get me wrong, I don’t think pre-registration solves all problems, but I just don’t see why it is likely to be unsuitable for long-term field studies. (Pre-registration doesn’t preclude additional exploratory work, not does it have to be about significance testing.)
The reason not to pre-register is that, while it might provide some slight additional margin to ensure an abberant result isn’t mis-confirmed by slipping through some statistical worm hole (forking paths etc), scientifically it’s pointless. Designing a sound analysis in which variables are controlled and/or covered by reliable assumptions – or even by less reliable assumptions for which the variation might be somewhat understood – is many times more beneficial than pre-registration.
The point is probably more relevant to ecology than the social sciences since in ecology the fundamental data is so much more reliable than in the social sciences. IOW: since ecology is built on a strong foundation of replicable science, it’s easier to construct a sound analysis. In social sciences, OTOH, lots of “experiments” are little more than guesses about social behavior in which people are attempting to confirm or reject some pulled-out-of-the-arse hypothesis with extremely poor meausrements. In that case, pre-reg is more likely to screen out bad results, since the proportion of bad results would be many times higher.
Oh – and by the way the biggest “statistical wormhole” isn’t statistical at all. It’s making incorrect assumptions about control of the variables. If you make a bunch of incorrect assumptions that you can ignore variables that in reality significantly influence the outcome of the experiment, then pre-register all you want, your result would have a high chance of being “confirmed”, since you’ve eliminated by declaration most of the realities that would cause it to be rejected.
Plugging again for ecology: ecologists, geologists, chemists, physcists etc are much more careful about making claims about ignoring this or that variable than social scientists (who often don’t even think about it). I’m not sure if that’s culture, or just the fact that, on that side of science, reality has a habit of emerging much more quickly and erroneous claims are thus exposed.
Responding to Chipmunk here as I can’t below:
Designing a sound analysis and pre-registration are not mutually exclusive. If both are beneficial in terms of the amount to which the rest of the world trusts our research, and the pre-reg cost is minimal or yields other benefits (as Andrew describes for simulating data), then why not do both?
I think that it’s a stretch to say that ecology is built on a strong foundation of replicable science. Ecology has abundant measurement error, proxy variables or concepts that don’t map on to the underlying reality, abundant issues of scale, unrepresentative sampling/unclear statistical populations, and a whole host of other issues. Not to mention things that cross the sciences, such as NHST, and people not understanding the difference between causal and predictive inference.
The issue is collecting the data in the first place, not analyzing some existing dataset. You have to scout a place, build it up, figure out what works and doesn’t, publishing as you go to keep the site funded. What equipment you can use, what samples you can return home, what species you’ll see, etc.. are all very stochastic and plans are often shaped as much by these forces as much as by the questions underlying the research. It’s inherently a process of “making do” and it’s rarely anywhere near perfect. At the same time, it’s not so hard to learn something new because often these places are poorly described over long timescales. With good observation shaping what to measure, signal just jumps out. The idea that this could be preregistered and provide any of the ostensible value it does elsewhere requires a loosening of the definition of preregistration so far that it’s little more than “keep a lab notebook”, which has been a part of biology and ecology for decades.
Pre-registration is pretty standard in the medical literature. Pleasingly, it seems to have got rid of p-value discontinuity at 0.05. See Decker and Ottaviani “Preregistration and the Credibility of Clinical Trials” Medrxiv 2023.