Clint Stober writes:
I would like to let you know about a paper my colleagues and I recently published in Perspectives on Psych Sci, and here is the preprint. We take a critical look at estimation accuracy across the behavioral sciences, using a hypothetical lab reporting random conclusions as a benchmark. We find that estimation accuracy can be so poor that it’s difficult to tell current practice apart from such a lab. It’s a short, but hopefully thought-provoking, paper that provides a different perspective on calibrating tests and the challenges of interpreting small effects. It certainly relates conceptually to Type S and M errors. Perhaps you and your readers will find it interesting. Links below to the article and the pre-print.
I’ve published in the journal Perspectives on Psychological Science, but more recently I was upset because the journal published a lie about me and refused to correct it. That said, journals can change, so I was willing to look at this new paper.
I like the idea of using “this idea of random conclusions to establish a baseline for interpreting effect size estimates.” This is related to what we call fake-data simulation or simulated-data experimentation.
It’s kinda what “hypothesis testing” should be: The goal is not to “reject the null hypothesis” or to find something “statistically significant” or to make a “discovery” or to get a “p-value” or a “Bayes factor”; it’s to understand the data from the perspective of an understandable baseline model. We already know the baseline model is false, and we’re not trying to “reject” it; we’re just using it as a baseline.
Isn’t this just essentially a type of null model, i.e. essentially what you would see if your data was all noise?
See here for an example in the psychology of art (Gaussian white noise as a null model of a time series): https://bactra.org/weblog/666.html
At least psychologists are beginning to learn.
It is *worse* than flipping a coin. Eg, here is spinal cord injury (~10% replication rate):
https://www.sciencedirect.com/journal/experimental-neurology/vol/233/issue/2
And here is cancer (~50% replicated, but only out of the ~50% where a replication was possible in principle):
https://www.nature.com/articles/d41586-021-03691-0
AFAICT they have continued on with the same NHST regardless. It won’t stop until people raise their standards to require replications and quantitative predictions for continued funding.
Hi Andrew!
First of all, I just want to say, I’ve been a big fan of yours for nearly 10 years now, ever since I was in college getting my BS in mathematics/statistics, and you’re knowledge and insights have played a major role in my development as an data analyst/poor-man’s-statistician over that time. Whenever I come across something I’ve been unsure about in relation to statistics/scientific research, your blog posts and publications are some of the first places I look to for information and considerations.
So, thank you for everything that you do; you’re truly making a wonderful difference in the world!!!
The main reason for my message today, though, is that I just came across [a preprint](https://osf.io/preprints/psyarxiv/7um9a), which is essentially an attempt to create a data-driven taxonomic system for psychological symptoms/syndromes to potentially, eventually, replace the DSM (it even seems to be an approvement over the HiTOP model), and I was hoping to get your thoughts on it!
In particular, I wanted to get your thoughts on their data collection/manipulation methods and, *especially*, their overall approach to the research project, like using the Open Science Framework (OSF), including as much detail as they did in both the main article and their supplementary materials about things like every single data manipulation/cleaning choice that they made, actually following a pre-determined data analysis plan, etc.
I haven’t read many recent papers outside of clinical trials research/statistics/data science in awhile, in part because it always felt a bit frustrating how barren most methodology/analysis sections often were, particularly when it came to their data cleaning and analytical choices, how reliant they were on p-values for just about everything, reproducibility/replicability problems, garden of forking paths issues, etc. *Especially* within the field of psychology, which seemed especially prone to such problems, for whatever reason.
But this paper… If there ever was a “right way” to conduct observational, survey-based, research studies, it feels like this group has actually done it!!
Maybe I’m just out of touch with how much better research papers have gotten in other fields over the past couple years, but it felt incredibly refreshing to read this paper!
This is the kind of paper I’ve always *wished* other papers were like and, unless I’m forgetting/overlooking something, I’d even go so far as to say that this is a superb example of how scientific research *should* be done (and shared), particularly for non-experimental research with a lot of inherent variability (e.g. human’s survey responses about their own thoughts/behaviors)!
Is it just me, or is this a particularly well done study/paper?
Or, is this sort of paper becoming more of the norm now?
Is there anything that comes to mind that they could have improved on?
Zack:
I took a look at the linked paper, and I don’t know what to say—it’s too far from any of my areas of expertise. I suppose that a statistical what-goes-with-what analysis of psychiatric symptoms can be valuable, as long as the results are not over-interpreted.
Dear Sir
Define the difference between “Making a discovery” and “understanding the data”
Yours Sincerely
Alan:
I don’t have any definition here; maybe it’s just a matter of scale and emphasis. To understand the data we need to make many micro-discoveries. If you look at the many applied examples in Bayesian Data Analysis, Applied Regression Using Multilevel/Hierarchical Models, and Regression and Other Stories, you’ll see lots of examples of “understanding the data” but none of “making a discovery.” Or, maybe they are discoveries, just very small ones.
I’m fairly confident that if you surveyed psychological researchers who are familiar with your writings on statistical methods in psychology over the last 10 to 12 years, the overwhelming majority would agree or strongly agree with a statement to the effect that your work implies that the field is inept and misguided. The fact that none of the reviewers or the editor took issue with the statement and refused to correct it says a lot about what people read into your underlying opinions of psychology. If you don’t want people propagating what you believe to be falsehoods about what your work does or does not “imply” then maybe try to understand why someone would arrive at that judgment and then adjust your approach to prevent further misunderstandings. Maybe take some time to consider why someone felt it was reasonable and accurate to claim your 2014 paper implied that you hold the view that the field of psychology is inept and misguided, and why people reviewing that statement tacitly agreed with the reasonableness and accuracy of that statement. They didn’t arrive at this belief in a vacuum. As the person who talks such a big game about how much you love criticism, it’s ironic how defensive you get when someone suggests that maybe you’ve been kind of an asshole to psychology.
Psychology *is* inept and misguided. He’s trying to save it from psychologists.
Sentinel:
No, I don’t think it’s ok for people to lie about what I wrote, just because they want to make a point or because they want me to take some time to consider something or whatever. If somebody wants to make a point or you want me to take some time to consider or whatever, I think they should do so directly, not by publishing lies about me. I’d actually be happier if they call me an asshole, as that’s a judgment of opinion which can’t be true or false.
And, yes, I love criticism (or, as you put it so charmingly, I “talk such a big game” about how much I like criticism). That doesn’t mean that I think it’s acceptable for people to lie about me.
This is a creative use of the word “lie.” The claim was not in quotation marks, so it doesn’t matter whether you used those words or not, the question is whether it is a valid summation of your arguments. And I think it is, or at least close enough that I wouldn’t have published your “correction” either.
Specifically, claiming that NHST is an invalid methodology is an implication that nearly the entire corpus of work in psychology is invalid. And there are numerous other examples like this where general practices in the field are trashed. The fact that you yourself have not rolled it all up into a general summation of your views about psychology? That’s why we have the word “implicit.”
Looking over your written work, it looks like you have anticipated this accusation and have tried to avoid having anyone be able to use a direct quote as confirmation, leaving space for you to feign indignation. It’s easy to find stuff like this (from the linked Gelman paper):
“The combination of high variation and small sample sizes in the [psychology] literature imply [sic] that published effect-size estimates may often be overestimated to the point of providing no guidance to true effect size.”
Sounds inept. Inserting “may often” does not help.
This is exactly the sort of circumstance Shakespeare had in mind when he wrote “methinks thou protestest too much.”
I’ll be happy to state it for the record, any field using NHST often enough to be considered a “dominant, or important” methodology is inept.
Of course I have little cache among social sciences researchers, with the possible exception of a few tens of people who read this blog regularly.
But this one shouldn’t be very controversial. NHST is not scientific. Even the frequentist die hards such as Mayo agree, and prefer that people test their actual hypothesis not straw man null hypotheses.
The whole point of “severity” is to check how stridently you’ve tested **your own hypothesis** not a straw man.
Matt:
I’m not “feigning indignation.” I’m actually indignant. Disagree with me all you like, but don’t tell me what I’m feeling.
The quote was, “some critics go beyond scientific argument and counterargument to imply that the entire field is inept and misguided,” and it cited a reference in which I did not do that.
Again, if people want to disagree with specific things that I’ve written, they should go for it.
And, no no no no no no no no no, to say (correctly) a lot of bad papers have been published and publicized in psychology is not the same as saying “the entire field is inept and misguided.” It’s just not.
Daniel:
There’s a lot of bad stuff published in psychology journals, and in medical journals, and in statistics journals, and everywhere else. There’s also a lot of good stuff. I agree that null hypothesis significance testing is often a bad idea; on the other hand, ultimately these are just methods that can be used well or poorly. Researchers can and do perform excellent research using null hypothesis significance testing. Other methods could work too, but that doesn’t mean the science is bad. Conversely, I’ve seen people use the sorts of Bayesian hierarchical models that I absolutely loooove, but in the service of bad science!
“I agree that null hypothesis significance testing is often a bad idea; on the other hand, ultimately these are just methods that can be used well or poorly.”
Dear Dr. Gelman,
In my limited experience, to the best of my understanding, I believe that even the very expression “null hypothesis significance testing” is incorrect. The word “significance” is an English synonym for “relevance” or “importance;” and as stated by Neyman and E. Pearson (NP) themselves, their “rule of behaviour” provides no information about the hypothesis itself or the individual statistical test. In other words, that concept of significance is certainly not statistical (i.e., it is not related to the relevance/importance of the statistical outcome) but is decision-inferential: “significance” is achieved only in the “long-run” (understood as “a high number of equivalent replications”) through multiple decisions disconnected from any specific statistical evidence of the single study, and is solely related to limiting type I errors to a predetermined frequency threshold α. Indeed, Neyman referred to α as the “level of [decisional] significance.”
Based on this, I personally find it nonsensical to speak of any “statistical significance” in the context of NP. For example, to state that “null P > α” equates to a “statistically non-significant” result is doubly absurd because 1) the result of a statistical test is much more than just the P-value for the null hypothesis, and 2) how can a statistical result be defined as “statistically non-significant” based on a decision rule that is mathematically precluded from providing information about the statistical test itself? Or more generally, how can a mere rule of behavior provide information on the “statistical significance” of a result?
Some might argue that NHST mixes NP’s “decisional significance” with “statistical significance” as a measure of refutational, statistical evidence: nonetheless, Greenland has shown substantial differences between decision and divergence P-values (especially concerning the so-called “equivalence” testing).
Therefore, I apologize in advance if my comment may seem drastic but, based on my current knowledge, I consider NHST a collection of fundamentally flawed ritualistic practices that would require redefinition from scratch to be properly understood (especially in the “soft” sciences). On this point, I would like to clarify that my critique of NHST is independent of the well-known problems of replication variability: Simply put, even assuming perfect replicative equivalence, I believe that even the name itself—as well as the expression “statistical significance”—has very little to do with the meaning of the procedure. Perhaps it could be called something like “Error Limitation Procedure” (ELP)? Or, ironically, HELP!