Reese Richardson reports on a recent study he did with Spencer Hong, Jennifer Byrne, and Luís Nunes Amaral, entitled “The entities enabling scientific fraud at scale are large, resilient, and growing rapidly.”
That’s a title that doesn’t mess around!
Richardson writes:
1. Editors abuse their positions of authority to collude with authors to publish problematic articles en masse. . . . certain authors seem to have a preference for having their articles handled by these flagged editors. For PLOS One, we identify a network in which these flagged editors were all handling each other’s submissions to the journal.
2. Networks of image duplication can be thousands of articles wide and these articles tend to appear in the same journals at around the same time. . . . articles connected by shared images also tend to get published at around the same time and in the same venue. This suggests that paper mills are capable of both producing articles and getting them published in a highly coordinated fashion–wholesale, not custom-made. Publishers often claim that paper mill products are published because they have slipped through the cracks. This vignette instead suggests a model where paper mills have relatively open pathways into journals–likely facilitated through the knowing cooperation of editors, as suggested in the first vignette.
3. Broker organizations are capable of placing articles in journals on demand and adapt well under adversarial conditions. . . .
5. Publishers understand that systematic fraud underlies the bulk of their integrity issues. We show that most retractions are now issued in batches, alongside ten or more retractions in the same journal on the same day. . . .
6. The integrity measures used to contain systematic scientific fraud are dwarfed in scale by the problem itself. We assemble a corpus of suspected paper mill products . . . this corpus has been growing in size (annual count doubling every 1.5 years) at a rate far eclipsing the growth rates of all scientific articles (doubling every 15 years) . . .
#5 here is particularly interesting to me. My head’s still stuck in the 2010-2015 Psychological Science era, a time when the absolute top journal in the field was routinely publishing junk science. At the peak of the problem, I think that more than half–no joke–of the papers in Psychological Science were crap, just pure combinations of noise mining, hype, and the sort of theory that was so flexible that it could explain any pattern or its opposite.
None of this was from paper mills, and I believe that very little of it was Wansink/Ariely-style fabrication or fraud. It was just bad science, done by highly-connected, highly-credentialed researchers who were working in a sort of unintentional parody of the scientific method. The sort of bad work that led to the terms HARKing, p-hacking, and questionable research practices. We discussed some such papers here, here, and here.
Based on this experience, I’ve been going around for years saying that the big problem in science is not fraud but rather well-intentioned bad work. Honesty and transparency are not enough.
But that was then, this is now. In this new era of paper mills and chatbots, where the marginal effort of writing and publishing a paper is essentially zero, there is more and more motivation for fraud. And, yes, typing some prompts into a chatbot and producing a paper is fraud, in the same way that publishing textbook excerpts as if it were new research is fraud, or copying from wikipedia as if it were new research is fraud, etc etc. It doesn’t require fake data and it doesn’t require some cackling Snidely Whiplash attitude. It can be some schlub sitting at a computer terminal who wants to get his contract extended or get admitted to a Ph.D. program or whose adviser is pressuring him to get some publications . . . But it’s fraudulent publication, not the same as bad research (which is actually research, it just happens to be useless because bad measurement and kangaroo).
I guess it still depends on the journal. But even if the top journals stay mostly immune from fake papers (so, for them, we can be more concerned about the bad research they publish), all this paper mill stuff has an impact, because it affects citation counts.
One option in evaluating research would be to follow the lead of economics, where I’ve been told that pretty much all that matters is the number of publications in “top 5” journals, and it doesn’t really matter how many citations you have or what you’ve published elsewhere. The downside of restricting to the top 5 is that this can enforce conformity (in writing style, research methods, subject matter, and conclusions), which leads to this sort of thing. As an interdisciplinary person, I like to publish in lots of different places.
The other weird thing is how we’re proceeding on multiple tracks.
On one hand, the current system of science is flying apart. Here’s Richardson:
We start our article by conceptualizing the scientific enterprise as one large public goods game . . . everyone makes some contribution to the pot and receives some reward realized through the collective sum of contributions . . . Scientists can earn long, rewarding careers. Private-sector firms can capitalize on new technologies. States collect more taxes from a healthier, wealthier workforce. . . .
Here’s what I’ve come to understand over the last three years . . . the scientific enterprise is now witness to widespread, organized defection from the scientific public goods game. Large swaths of players, among them many scientists, reviewers, editors and publishers, are choosing to no longer make genuine contributions to the pot. Parascientific organizations (like ARDA, paper mills and the groups of collaborating editors) now facilitate and profit from mass-scale defection. Many scientists, especially in countries where the resources for doing genuine science are more scarce, are now trained in contexts where defection is the normative behavior.
He continues:
Some model public goods games integrate a mechanism by which defectors are punished. While this can be effective at mitigating defection under certain circumstances, our study shows that the punitive measures employed to enforce science integrity, like retractions and de-indexing, are currently applied far too infrequently to meaningfully increase the costs of defection.
In the meantime, the resources and incentives to doing good science are declining:
Meanwhile, the United States government is dismantling research support infrastructure and funding wholesale, defecting from their role in our immense public goods game and ensuring that contributions by taxpayers will also wither. In effect, competition for resources will only grow more fierce and it will become more and more difficult to make genuine contributions. . . .
If the model public goods game offers any prognostication, it’s that the current paradigm, where defection is the winning strategy, ensures that genuine contributions will only decay from here. We will all be worse off for it. Anyone that has studied industrialized scientific fraud has seen the future of our scientific enterprise . . .
I have seen the future of science. It is ruled by bitter competition instead of collaboration, pageantry instead of exploration. Bright minds beginning careers in science will be taught to debase their training for drudgerous pursuit of meaningless metrics. Those willing to toil over genuine questions will necessarily lose out to those that can furnish cheap answers.
Many scientists worldwide already inhabit this reality. So, we all may. If this vision comes to pass, humanity will lose its most potent engine for progress and its most abundant source of wonder.
So, yeah.
On the other hand, when doing my own writing and research, I’m still living in the world I grew up in, where we craft our articles one at a time and shepherd each through the reviewing process. I keep doing this. This is weird. I don’t know what to think.
Grim indeed. And AI is only beginning to have effects – which are surely going to exacerbate the problems. Most of my experience has been an non R1 universities so I won’t claim to characterize practices there. But for the other 90% of universities, I have ample experience. My observation is that academics are reluctant to attempt to judge the quality of colleagues (either job applicants or current colleagues) research – instead they defer to number of peer-reviewed publications (which often numbers very few). I believe this reluctance comes from both an avoidance of personal conflict as well as a fear of exposing oneself by criticizing another’s work (and possibly being wrong, or more importantly, being subject to the same critique).
Structurally the problem seems serious to me and resistant to solutions. Ultimately I believe it is best addressed by academics taking each other’s work seriously – whether it be research or teaching. I think the quality impediments applying to teaching are even more serious than those for research.
I have wondered a lot about how I am able to tell or know that what I am thinking when judging a scientist’s work is correct, or at least appropriate and valid enough to note. It might be very hard to tell or know this. Concerning my most recent manuscript I have wondered whether I am seeing things correctly, or whether I am making mistakes in my reasoning or judgement, a lot.
I find it very hard to determine whether one is viewing things correctly, and in my particular case I haven’t got any colleagues or peer reviewers to also take a look. So, it comes down to your own judgement about your own ability to spot possible mistakes or strange things in what you read. In some cases, this process is assisted in a way by coming across criticism by others concerning the same paper or study, which has been helpful for me, but that is not always available.
There is nothing new under the sun. Even when science was a small pool, hasn’t this always been the case (“..bitter competition instead of collaboration, pageantry instead of exploration. Bright minds beginning careers in science will be taught to debase their training for drudgerous pursuit of meaningless metrics”)?
See:
Einstein failing to secure a professorship after his miracle year. Mendel almost being forgotten. Boltzmann being driven to suicide…
I am naively optimistic that this will all work itself out. There are parallels to looking at democracy up-close and at short time scales, it is always horrifying to see how the sausage is made.
Student:
I’ve always thought that an important characteristic of a good scientist is the capacity to be upset, to recognize anomalies for what they are, and to track them down and figure out what in our understanding is lacking. This sort of unsettled-ness—an unwillingness to sweep concerns under the rug, a scrupulousness about acknowledging one’s uncertainty—is, I would argue, particularly important for a statistician.
But the capacity to be upset is not so common. Intellectual complacency was a problem in the past, too! So, yeah, maybe you’re right that these problems have always been around.
I agree – I am upset as well. Hope & timelessness is my kneejerk reaction to doomerism. I kept reading for a kernel of hope out in the above text and kept getting relentless despair. I don’t think all will be lost.
What’s new, I think, are the metrics. Back in the day (<~1975), science didn't suffer from Goodhart's law, because there were few objective metrics that could be gamed. Of course, subjective evaluation has its own problems …
Ziggy:
My crude understanding of the history goes something like this:
Before 1960: Academia was small. People pretty much just hired their friends.
1960-1980: Academia was growing and it was hard to fill all the slots. Pretty much anyone with a pulse was getting hired.
1980-2000: Jobs are getting tight and now there are tons of Ph.D.s, so people go back to hiring based on connections.
2000-present: Metrics!
The most interesting element in the Richardson et al. study is the role of editors. Naively, I assumed they were just fallible gatekeepers, selected for markers of expertise and with some leeway for scrupulousness. Depending on how much effort they wanted to put into the job, they could be bloodhounds for shoddy work or dozing night watchmen. But this piece says it’s much worse than that — they are collaborators, knowingly enabling industrial scale slop.
(a) Do we believe this? (b) If it’s true, what does that say about the selection and assessment of editors? The problem as claimed in this paper is so vast that it must have infected all levels of research and publishing.
This is from my most recent manuscript, in which I mention some research concerning the possibility that journals, editors, and peer review might not be beneficial. I actually saved the Richardson et al. paper for that manuscript, but ended up not using it. Here’s a section I wrote that includes some findings involving editors:
“The investigative report concerning the fraud case mentioned earlier (Levelt et al., 2012) mentions that: “Virtually nothing of all the impossibilities, peculiarities and sloppiness mentioned in this report was observed by all these local, national and international members of the field, and no suspicion of fraud whatsoever arose.” (Levelt et al., 2012, p. 53). Perhaps even more noteworthy is that co-authors of the fraudulent psychological scientist reported more than once that editors and reviewers actually encouraged irregular practices (Levelt et al., 2012, p. 53). In line with these findings of the investigative report, it can be argued that journals, editors, and peer reviewers in general have contributed to certain problems (e.g. researchers engaging in QRPs) by selecting papers based on novelty, flow of the narrative,
and perfection (see Giner-Sorolla, 2012, p. 567; Levelt et al., 2012, p. 53; Schimmack, 2012, p. 563). Next to this, it has been reported that editors have guided authors to add citations from their journal in order to inflate the impact factor of their journal (see Fong & Wilhite, 2017). This “coercive citation” can even lead to authors behaving strategically by adding references that have recently been published in the journal they are submitting their work to (see Chorus & Waltman, 2016, p. 7; Fong & Wilhite, 2017, p. 13). The peer review process can lead to authors thinking and behaving strategically in other ways, for instance by adding references authored by potential reviewers, and by avoiding criticizing the work of possible reviewers (see Binswanger, 2014, p. 57). Such things might be done by some authors because the peer review process may in certain cases not be truly anonymous (see Binswanger, 2014, p. 58), and it can involve a small group of people who influence what gets published (see Anderson et al., 2007, p. 452; Heesen & Bright, 2021, p. 647; Tennant & Ross-Hellauer, 2020, pp. 3-4).”
This rejecting “chance” idea (along with replacing replications with peer review) IS the institutionalized “fraud” that not just enables, but incentivizes, these other problems.
I put “fraud” in scarequotes because most of those implicated don’t even realize they are doing anything wrong due to overreliance on argument from authority/consensus heuristics.
Just do direct replications and compare the predictions of your theory against new data. It is really simple in concept, but much more time-consuming and expensive than the “fraud”. As a bonus, this will automatically catch all deliberate frauds as well.
This seems to be part of the general enshittification of public/private interactions in our modern world and the unwitting (maybe not so unwitting) creation of incentives for bad actions. Still, one doesn’t need to wallow in the garbage!
I’m trying to understand why the Richardson et al PNAS paper seems quite foreign to me. I think it’s due to its wallowing in garbage. They set out to identify sordid practices in scientific publishing and they found it in abundance in the disreputable fringes of the publishing industry.
For example they searched for journals with high numbers of apparent “paper mill” products and present the “top” 50 of these (Supplement Fig S16). Many (not all) of these are garbage journals. Scientific Reports is a decent journal IMO (I‘ve published there) but it’s in their top 50 for suspected paper mills – the number of suspected paper mill products (132) covers a period in which around 115,000 papers were published – so around 1 in 800 papers in their covered period is suspect. Is that such a big deal? On the other hand Int. J. Biol. Macromol. has more like 2% of published papers as suspected paper mill products. But Int. J. Macromol. is pretty widely recognised as a journal of negligible value – I would never publish there, and have never cited a paper in it. Likewise with PLoS One which the Richardson paper makes a large study of – it’s one of those journal of last resort for when you’re struggling to get a paper somewhere good – I would never publish there either (that’s not to say that some decent papers aren’t published in PLoS One). Likewise everyone knows that Hindawi journals are generally garbage, same with ARDA.
I think it’s important to recognise the context in which this highlighting of dismal practices in publishing occurs. Absolutely there is an abundance of garbage of the sort highlighted by Richardson et al. But it’s mostly recognisable garbage that lies in the massive rump in the ignorable depths of published “science”. If I wanted to know what some of these practices mean for the reality of publicly-funded science in the US for example, I would want to collate NIH-funded biomedical research (say) and assess the journals in which this research was published, and whether any of it fell into the “suspected paper mill” category. I expect almost none of it would be published in Hindawi or ARDA journals or Int. J. Biol. Macromol. or Benha J. App. Sci. or most of the other journals discovered in the Richardson trawl. If it was, it’s unlikely that NIH would be enthusiastic about continuing funding.
IMO there are lots of reasons why Richardson’s lament “Those willing to toil over genuine questions will necessarily lose out to those that can furnish cheap answers.” won’t come to pass at least so long as we live in societies that haven’t been taken over by the barbarians.
didn’t mean to write such a long essay.
shorter version: Just because you find lots of garbage in lots of dubious places doesn’t mean there’s so much garbage in the places that might actually matter.
I do agree with the final paragraph of the Richardson PNAS paper about the problems around the inability of AI to distinguish junk from quality. That’s surely not an insurmountable problem though. Perhaps one of the positve outcomes will be a realization of an objective and implementable set of criteria of quality.
“One option in evaluating research would be to follow the lead of economics, where I’ve been told that pretty much all that matters is the number of publications in “top 5” journals, and it doesn’t really matter how many citations you have or what you’ve published elsewhere”
For a field that is wrong often, I’m amused by this. But then they will have the chance to model why the top 5 journals are more profitable.
Hard to square this kind of doomerism with what is happening in, say, Cosmology (the JWST in particular). It seems to me that there are fields of science that still perform quite well. Whereas the stuff that’s always been sketchy (psychology, Evo Devo, economics) has gotten sketchier, with the process of mathematization in these fields not having made them more rigorous.
How does one determine what are the top journals? I come out of psychology. A quick check of the American Psychology Association’s website shows 56 divisions.
A Division 14 (Society for Industrial and Organizational Psychology) member is unlikely to have much or any interest in anything from Division 40 (Society for Clinical Neuropsychology) researchers–well unless they are dealing with some senior executives.
I still can’t get over the headline interpretation from Reese Richardson’s blog that fronts this thread but perhaps I’m just validating a very successful example of clickbait! Still, it seems unfortunate to present something that is blatantly dodgy, and it’s a little disappointing coming from a scientist.
IMO Richardson has chosen to equate seperate phenomena that don’t have much overlap. There’s no question that the disreputable practices he documents in his perfectly decent PNAS paper have and are occurring. But the “Future of science…” handwringing which insinuates the spread of these practices to all science??
1. Publications: People might have been put off by the use of “garbage” in my post above, but many of the journals in Richardson’s PNAS paper are truly dismal. Their Supplement figure S16 is one where you can see which journals dominate one of the problems – suspected paper mill infiltration in this case. There’s a bunch of Hindawi journals (truly dodgy) and three of the top five (all Hindawi’s!) have been shut down. Incidentally, that indicates some of the handwringing is misplaced. One can compare this list of journals with the journals that scientists funded by the NIH (say) publish in. This info is available but not easily I think, but it’s easy to find a list with numbers of publications from the NIH itself in the calendar year 2025 (Google “Nature Index, National Institutes of Health”). Research funded by the NIH isn’t published in crap journals and it’s not clear why anyone would think this is going to change in a doomsday scenario.
2. Location: A high proportion of the issues highlighted by Richardson et al come from outside Europe, N America, Oceania, Japan, Korea, Singapore et al. That may sound elitist but it’s obvious, for example, that China especially has been a problem. 75% of the “top” 100 Universities for retractions are Chinese; returning to our favourite Hindawi journal, 9,600 articles were retracted in 2023 – over 8000 involved Chinese scientists.
3. Funding: An obvious reason why Richardson’s gloomy prognosis is unlikely is funding. NIH funds biomed research at around $45 billion per year, charities/philanthropic ~ $30. Funders, including industrial funders, aren’t going to piss away vast sums of money without feeling like they’re getting a return. In fact, China is starting to take scientific misconduct more seriously and they’ve recently introduced a policy of targeting Uni’s and research institutions if these don’t properly investigate researcher misconduct; it wouldn’t surprise me if the future progresses in the opposite direction from Richardson’s. You only have to go back to Richardson et al.’s list of 50 top “paper mill” journals to see that over half of these have been shut down or delisted by publication archives or put on “warning lists” so these are not insurmountable issues.
OK so I’ve been properly caught by a successful clickbait! But this annoys me – there seems to be an appetite for trashing science in the US; two of the more prominent “data sleuths” have highlighted the issues of extrapolating from specific instances into “science in crisis” narratives (Dorothy Bishop) and allowing critique of science to become “weaponized” against science and scientists (Elisabeth Bik).
Theres little difference in scientific return between the ~10-30% replication rates seen in the spinal cord + cancer reproducibility projects vs ai slop. Both are negative, generating misinformation with very low chance of a winning lottery ticket.
So we need to consider what return the funders are actually looking for. Selling a product for industry, creating jobs for the government, getting donations for philanthropic organizations.
All are only indirectly (and unnecessarily) correlated with scientific returns.
From Table 12 of the supplementary material of Wilhite and Fong (2012) “Coercive citation in academic publishing”
“Journals identified as coercers by survey respondents. Number of coercive observations represents the number of times a journal was identified by independent survey respondents as requesting self citations that (i) give no indication that the manuscript was lacking in attribution, (ii) make no suggestion as to specific articles, authors, or a body of work requiring review, and (iii) only guide authors to add citations from the editor’s journal.”
Here are the first five journals named:
Journal of Business Research
Journal of Retailing
Marketing Science
Journal of Banking and Finance
Information and Management
Economics focus on Top 5 is intellectually damaging, IMO. There are easily another 25 journals, maybe 50 (good general journals and top field journals) with high standards (at least as measured by rejection rates). But once you go below that level, publications do you more harm than good.
And economists were using “data mining” as a pejorative decades ago.