A researcher who wishes to remain anonymous writes in with a question:
I am reaching out to you regarding a weird phenomenon taking over the environmental science literature.
The field has been flooded by papers that appear to me to be from a paper mill. The papers come from dozens of different researchers primarily not in North America or Europe, but they generally follow the same template of (i) downloading time series data from the World Bank and (ii) applying time series regression to infer causal links.
The list of papers is in the thousands, but here is a sample of several hundred papers that refer specifically to autoregressive distributed lag models (ARDL) in the title or abstract. This subset is predominantly from one journal, Environmental Science and Pollution, but a wider search includes many other journals from multiple publishers.
I am part of an informal group that investigates systematic research fraud in our free time. Our experience has been that the first step should be to identify scientific errors in the work. Unfortunately, other evidence of misconduct (like coordination, manipulating peer review, fake authorship etc.) is usually only circumstantial, so authors have plausible deniability. However, I am an ecologist, not a statistician, which is why I am contacting you.
I know enough to understand that there is no single trick to causal inference, and that statistical interpretation is exactly that: an interpretation. But I hope you might (or could recommend someone who might) help identify a statistical “red flag” in the way these papers are applying regression models? Is there any heuristic we might use to identify wholly inappropriate stats and flag questionable papers? It would be helpful to distinguish between nonsense statistics vs. valid statistics applied improperly.
Unfortunately, no, I don’t know of any heuristic, beyond using tools such as GRIM to check for consistency of reported numerical results. (That’s how they tripped up the Pizzagate guy.) Unfortunately, I think the scientific publication process is out of control and will only be getting worse.
Here’s an easy one: it’s in NBER, it’s wrong.
LOL. Except for this one:
https://www.nber.org/system/files/working_papers/w21143/w21143.pdf
Above a Swamp: A Theory of High-Quality Scientific Production
Bralind Kiri, Nicola Lacetera, and Lorenzo Zirulia
We elaborate a model of the incentives of scientists to perform activities of control and criticism
when these activities, just like the production of novel findings, are costly, and we study the
strategic interaction between these incentives. We then use the model to assess policies meant to
enhance the reliability of scientific knowledge. We show that a certain fraction of low-quality
science characterizes all the equilibria in the basic model. In fact, the absence of detected low-
quality research can be interpreted as the lack of verification activities and thus as a potential
limitation to the reliability of a field. Incentivizing incremental research and verification activities
improves the expected quality of research; this effect, however, is contrasted by the incentives to
free ride on performing verification if many scientists are involved, and may discourage scientists
to undertake new research in the first place. Finally, softening incentives to publish does not
enhance quality, although it increases the fraction of detected low-quality papers. We also
advance empirical predictions and discuss the insights for firms and investors as they “scout” the
scientific landscape.
(Subject to the irony that it has probably been inadequately verified, of course.)
Is the issue that (i) there are methodological errors in these papers, or that (ii) it’s a flurry of mediocre papers, simply data mining for correlations and not contributing anything to our understanding? If (ii), these papers are hardly alone. Less than two weeks ago, we saw that this could get you into the New York Times (https://statmodeling.stat.columbia.edu/2022/08/15/coffee-study-lower-dying-risk-html/#comment-2071635), which seems worse than than some poor fools padding their CVs.
To be only slightly…ahem….ARCH, is there any result in the entire autoregressive distributed lag literature not directly tied to a known direct physical process that *isn’t* just data mining?
I have a list of statistical “red flags” at https://journal.sjdm.org/stat.htm. These are more specific to psychology, economics, and (sadly) medicine. The list does not mention anything as fancy as time-series analysis, or anything like the dredging that you describe. Moreover, “red flag” does not mean “wrong”. For editors (me) and authors it just calls for careful attention and suitably qualified reporting in the paper. But, yes, often these are real mistakes. They would lead me to desk-reject a non-trivial proportion of papers in the most “impactful” psychology journals.
“Significance tests are most useful for well-controlled experiments, which are usually designed, often with great care, to make the null hypothesis essentially true when the experimental manipulation has no effect. ”
Excellent! Most excellent. This can’t be said enough. Every method has a specific use case. If people restrict the method to the use case, SURPRISE! it works! If they apply it to wildly different cases – SURPRISE – it will NEVER WORK! :)
“Note that “not significant” is not the same as “no effect”.”
But do you agree that “not significant” = no *detectable* effect? IOW, for a treatment of some fixed amount to a specified condition, “not significant” means there is no *known* effect at that level of treatment. I find it irritating when people lean on the direction of a “not significant” effect to claim that, in principle, the treatment is working / not working as expected. AFAIK, this is not true. If the effect of the treatment is not distinguishable from a random distribution, the apparent direction of the treatment is meaningless.
> effect of the treatment is not distinguishable from a random distribution
Well some treatment effects are more or less compatible with the data and assumed model but confidences and credibilities (posterior probabilities) only relate to the possible world defined by the assumptions of the model, not the actual world that produced the data. (perhaps google Greenland compatibility intervals)
A physics inspired logic which encourages thinking of models as being literally about the world discourages this distinction?
Significance tests are bad and wrong even for well-controlled experiments
Perhaps you can explain the thought process. Do you imagine the confidence/credible interval expands to infinity if it crosses zero? Or is it that the underlying likelihood/posterior (summarized by the interval) flattens out to be one everywhere within the arbitrary interval and zero everywhere else?
I really don’t get what people are thinking, but it appears to be common to believe you can’t say anything just because an arbitrary interval crosses zero.
Statistical significance is interpreted as a license for discovery.
Every time I try to get practitioners to explain the logic they quickly get confused and simply point to their training and to standard practice.
In other words, the only logic is historical circumstance.
Unfortunately we appear to have failed to elicit an explanation once again.
A frequentist hypothesis test will “work” if it happens to agree with the Bayesian analysis that is appropriate for the problem. Of course, you can’t know this until you do the Bayesian analysis, so you might as well just do that and skip the frequentist stuff.
Ha! Hilarious responses! :) Thanks to all of you.
I thought the funniest part was “Every method has a specific use case.”
Do these papers include any actual data analysis? Are the data plotted appropriately and checked for errors and outliers? Are they re-expressed when needed? Are the residuals examined? Any regression diagnoses? And do the conclusions discuss any meaning drawn from the analysis. All of these steps require human participation by someone who understands the discipline from which the data are drawn. Absent these steps, papers should be rejected anyway. But I’d expect them to be missing in work from a paper mill because they require time and effort.
Paul:
Good point regarding effort. That said, we’ve seen lots of bad applied papers using econometric methods where the authors put in lots of effort but still end up with unsupportable conclusions (for example the paper discussed here). Indeed, often it seems that the large amount of effort expended is taken as evidence that the conclusions are correct (this is sometimes referred to as “robustness”).
I skimmed your article. One factor not mentioned is term limits https://www.termlimits.com/which-states-have-term-limits-on-governor/ — 36 states limit the number of times one can win an election, but losers can run as many times as they wish. And narrow losers have more reason to run again. Another possible factor is ageism; voters may be disinclined to vote for elderly candidates.
Those aside, if one is willing to spend additional effort, why not expend it attempting constructive replication? What about elections for other offices? Is there any reason to suppose that there is something special about gubernatorial elections, vs. other state offices? Or federal ones? Or Members of the Legislative Assembly of Canadian provinces? Surely there are numerous sets of out-of-sample data against which the purported effect could be tested.
Ken:
The difficulty here is mathematical, similar to what is discussed in our paper on small effects: For any reasonable effect size, you’d need to have a huge sample size to detect anything. We also call this the kangaroo problem. A related problem is that any effect will vary over time and across contexts; it will be positive in some settings and negative in others.
Realistically, I don’t think the studies that you propose, looking at legislative assemblymembers, etc, will work. It will be difficult to put together the data (that was an issue in the dataset with winners and losers in governors’ races) and the estimates will be too noisy compared to any realistic persistent effect size.
I think there are more interesting things to study in political science. One reason I wrote about the topic at all is because I think these sorts of simplistic, push-button notions of causality (win the election and you live 5 to 10 years longer) are a problem in much of social science thinking, and it can be helpful to understand the reasons why these sorts of studies fail and the mechanisms by which researchers confuse themselves.
A quote from one of the papers (FU Khan et al.):
“The counterfeit alterations in regressors and their effects on regression have been revealed in this research.”
At least they’re frank about it!
from the journal site: as of August 2022, they are on issue 39 and page 59947 (sic) for the year…so yeah, it’s likely that “a couple” things slipped through the cracks
https://link.springer.com/journal/11356/volumes-and-issues
“Unfortunately, I think the scientific publication process is out of control and will only be getting worse.”
I think this is the key thing here. The env science/ecology literature is full of people churning out dull or pointless papers because they can get hold of open data and have a method that is broadly “accepted” (how many hundreds of papers are there based on people creating Species Distribution Models using MaxEnt from GBIF data?)
I don’t think it’s necessary to invoke fraud or a paper mill. An economic/prestige incentive and a lack of serious critical engagement with one’s subject is more than enough. (I also wonder whether there is a tendency for people to suspect the worst (i.e. conscious fraud) when the authors are non-Western and the standard of English is not as polished as it might be. It’s easier to convince people that your work is useful when you have a sophisticated writing style.)
You make a good point about the being careful not to invoke paper mills when something could just as reasonably be explained by shallow bandwagon research. But there are other reasons I assumed the worst (FYI – I posed the original question to Andrew). For example, the sheer number of papers, which seem to be doubling each year. The unrealistic number of citations to manay of of these papers point toward citation manipulation (e.g., should we reasonably expect that a paper on renewable enegery in Turkey will receive >280 since it was published in 2020; 28 times the journal impact factor?). It seems strange to me that collaborators from a handful of countries in Asia and the Middle East will collaborate on papers about unrelated third countries? Of course, none of these features on their own is evidence of fraud or misconduct, but viewed together they do raise eyebrows…
…That said, none of it matters if the studies are obviously statistically incorrect. An incorrect paper is an incorrect paper, regardless of how it was written. But I lack the statistical expertise to quickly judge the correctness, hence my qestion.
Regarding your comparison to SDMs using MaxEnt models and GBIF, I’m not sure the comparison is a good one. I agree that there are too many superficial SDM papers out there, but a more accurate comparison would be to a set of authors that seem to produce dozens of MaxEnt papers using what appears to be a template, each time changing the focal species or geographic context…I may be wrong, but I haven’t seen anything like this yet. Holding thumbs I never will.
Lastly, you raise a good point about academic snobbery and a tendency to expect the worst from non-Western authors. This is a real cognitive bias. However, I would assume that such a bias would lead to *fewer* of these papers being published because reviewers and editors would be (unfairly) skeptical. We see the opposite: the papers are almost exclusively published by authors we would expect to experience unfair biases against them.
Thanks for the response. That all seems fair enough — I made little effort to dig into the papers except to look at a handful of the authors in ResearchGate.
My only quibble would be on the last point: I don’t see that these authors would necessarily come up editor/reviewer prejudice of they are publishing in journals dominated by their regional peers/collaborators/sympathisers etc. I suppose this would be the difference between a loose “citation cartel” and an organised paper mill.
Another point is perhaps that it’s easier to discern shallow copycat research when it’s done by authors publishing outside of their native language. So, with the SDM example, perhaps some authors are better at disguising vacuity, and so these non-Western examples are more easily spotted. But, as I say, I have made no effort to look into this. Good luck with your project.
Try this. Run (drive!) papers through a “new quantum hardware technology called Entropy Quantum Computing (EQC)”. If they don’t crash in your environment – publish.
Six seconds to solve “complex problem consisting of 3,854 variables and over 500 constraints.”
I’d say you’ll be able to come up with a lot less sensors – “statistical red flag” – but probably more constraints. Is it sufficient?
Looks like less than 3,854 variables and over 500 constraints.
“Sufficient statistic
https://en.wikipedia.org/wiki/Sufficient_statistic
Plus “Common variables associated with social science research”
https://www.researchgate.net/figure/Common-variables-associated-with-social-science-research_tbl1_303664446
And I’ll bet you’d love to get your hands on a “new quantum hardware technology called Entropy Quantum Computing (EQC)”.
”3,854-Variable Problem Solved in Six Minutes With Quantum Computing
By Francisco Pires
28 July 22
…
“QCI Achieves a Quantum Landmark for BMW by Solving 3,854-Variable Problem in Six Minutes
“QCI Using its Entropy Quantum Computing (EQC) System Achieved Superior Results in BMW Sensor Optimization Challenge
“LEESBURG, Va., July 20, 2021 –Quantum Computing Inc. (QCI) (NASDAQ: QUBT), a leader in accessible quantum computing, today announced that it has solved an optimization problem with over 3,800 variables in six minutes, delivering a superior and feasible solution. The Company achieved this landmark by applying a new quantum hardware technology called Entropy Quantum Computing (EQC) to the BMW Vehicle Sensor Placement challenge, a complex problem consisting of 3,854 variables and over 500 constraints. In comparison, today’s Noisy Intermediate Scale Quantum (NISQ) computers can process approximately 127 variables for a problem of similar complexity.”
…
https://www.quantumcomputinginc.com/press-releases/qci-bmw/
How about training an AI engine on your sample?
Background: I worked at IEEE during the fake paper escapades of 2014[1] which has seemed to continued into 2021[2]. I have also a patent in extracting concepts from scientific texts[3]. Hell – when I was at IEEE I found a fake, humorous article about the inner workings of an erotic massage tool – complete with graphic diagrams :)
Take a look at Tortured phrases: A dubious writing style emerging in science: Evidence of critical issues affecting established journals[4]. It talks a set of “tortured phrases” found in publications along with their correct, more commonly used phrase. For example, “big data” turns into (enormous | huge | immense | colossal) information. Artificial Intelligence (AI) turns into (counterfeit | human-made) consciousness. “Remaining energy” turns into “leftover vitality remaining energy” (sort of sounds like a Cialis commercial).
Let’s look at a few good ones from the article on arXiv:
“counterfeit neural organization” instead of “artificial neural network”: https://bit.ly/lol-ieee-1
“profound neural organization” instead of “deep neural network”: https://bit.ly/ieee-profound
“focal preparing unit” instead of “central processing unit (CPU)”: https://bit.ly/come-on-its-a-cpu
A quote from that last one: “The Raspberry pi board behaves like a focal preparing unit for observing and controlling gadgets joined to it.”
The sad thing is this – they do not care. These publishers worry about one thing – the big number above their search box of how many PDFs they have – as if they’re McDonalds in the 1980s counting the number of cheeseburgers sold. How many do you think are obituaries?[5] Or Table of Contents?[6] Or Magazine Covers?[7]
Academic publishing is a shameful industry – but these guys really milk it for what it’s worth.
The stuff above is the low-hanging fruit. Where you can really start looking at suspect publications is in an analysis of the citation graph. It’s exactly how Google determined which pages were garbage which should rank in the beginning.
—
Tom
linkedin.com/in/tomgriffin
t.me/tomgriffin
[email protected]
[1] https://www.nature.com/articles/nature.2014.14763
[2] https://www.nature.com/articles/d41586-021-01436-7
[3] https://patents.google.com/patent/US9323829B2/en
[4] arXiv:2107.06751 [cs.DL]
[5] https://ieeexplore.ieee.org/search/searchresult.jsp?newsearch=true&queryText=obituary
[6] https://ieeexplore.ieee.org/search/searchresult.jsp?action=search&newsearch=true&matchBoolean=true&queryText=(%22Document%20Title%22:%22Table%20of%20Contents%22)
[7] https://ieeexplore.ieee.org/search/searchresult.jsp?action=search&matchBoolean=true&queryText=(%22Document%20Title%22:%22Cover%22)&highlight=true&returnFacets=ALL&returnType=SEARCH&matchPubs=true&refinements=ContentType:Magazines
yes, interesting, thanks. even just looking through some articles randomly in Environmental Science and Pollution the other day, i caught weird phrases like that. at first you think, it might be ESL issues, but more likely it’s clumsy plagiarism evasion.
Yeah, I agree with all you wrote above. I wondered if the sentence flagged by a commenter above is an example of a tortured phrase: “The counterfeit alterations in regressors and their effects on regression have been revealed in this research.”
Can anyone on here come up with a valid statistical synonym that could be twisted into “counterfiet alterations” (my guess would be “counterfactual transformation”, but I am not sure that would be statistically meaningful in this context).
> Where you can really start looking at suspect publications is in an analysis of the citation graph. It’s exactly how Google determined which pages were garbage which should rank in the beginning.
I agree 100%. That is how this was brought to my attention (FYI, I worte the initial qestion to Andrew). So many of these papers are top-ranked in their field by citations (Web of Science Hot Papers, or Highly Cited Papers). It is clear that the majority of citations are to general sentences, not specific feature of the papers being cited…the problem with this is that it is nearly impossible to point fingers at the authors of the highly cited papers citation cartels are very sophisticated and avoid self-citation and quid-pro-quo citations between pairs of authors. Even if there is clear and obvious evidence of citation manipulation, how would one even deal with it? Like I said, journals won’t retract a paper becasue it was cited inappropriately by other authors in differnt journals. Similarly, the citing papers won’t face sanction, becasue no journal will act against a paper where, say, 5 out of >50 citations are inappropriate and out of context.
Long story short, I am unsure whether these papers on ARDL are the work of a sophisticated papermill, or simply a collection of independent authors who mutually benefit from keeping the research bandwagon going. It won’t matter either way if the underlying statitics are incorrect, which was why I posed my orignial question to Andrew.
Any paper with more than 10 Authors is useless, and possibly fake.
Any paper with 100+ “Authors” is garbage.
Anon:
But a paper with exactly 10 authors can be excellent.
Is there any way you could put up an accessible (i.e. free) version of the GRIM article? We non-university folks have to actually pay for articles.
Among other things, I would like to try the tests on my own papers. I know that there is nothing dishonest there, but I also know that I can and do make mistakes.
William:
The wikipedia article on GRIM has a link to a free version of the article.