“On the Past and Future of Null Hypothesis Significance Testing”

I happened to come across this cool article from 2001 by Daniel Robinson and Howard Wainer. It was published in the Journal of Wildlife Management, of all places—I didn’t know that Howard ever worked in that area!—but you can find a free version here at the Research Reports series of the Educational Testing Service.

The article starts with a bang:

In the almost 300 years since its introduction by Arbuthnot (1710), null hypothesis significance testing (NHST) has become an important tool for working scientists. In the early 20th century, the founders of modern statistics (R. A. Fisher, Jerzy Neyman, and Egon Pearson) showed how to apply this tool in widely varying circumstances, often in agriculture, that were almost all very far afield from Dr. Arbuthnot’s noble attempt to prove the existence of God. Cox (1977) termed Fisher’s procedure “significance testing” to differentiate it from Neyman and Pearson’s “hypothesis testing.” He drew distinctions between the two ideas, but those distinctions are sufficiently fine that modern users lose little if they ignore them.

They continue with some background:

Fisher understood science as a continuous process and viewed NHST in that context. He often used NHST to test the potential usefulness of agricultural innovations. He understood that science begins with small-scale studies designed to discover phenomena. Small-scale studies typically do not have the power to yield results of unquestioned significance. Moreover, Fisher recognized that the cost of getting rid of a false positive was small in comparison to the cost of missing something that was potentially useful. He knew that if someone incorrectly found that some sort of innovation improved yields, others would quickly try to replicate it. If replication repeatedly failed, the innovation would be dismissed.

Fisher (1926, p. 504) adopted a generous α of 0.05 to screen for potentially useful innovations “and ignore entirely all results which fail to reach that level. A scientific fact should be regarded as experimentally established only if a properly designed experiment rarely fails to give this level of significance.”

Think about that for a moment. The idea here is not that statistical significance implies “discovery”; rather, the statistical significance filter was to be used to decide that certain experiments (the ones that could not even reject the null hypothesis) were so weak as to be essentially useless in themselves.

They continue with more along these lines:

Fisher (1929) went on to say . . . “It is common practice to judge a result significant, if it is of such a magnitude that it would have been produced by chance not more frequently than once in twenty trials. This is an arbitrary, but convenient, level of significance for the practical investigator, but it does not mean that he allows himself to be deceived once every twenty experiments. The test of significance only tells him what to ignore, namely all experiments in which significant results are not obtained. He should only claim that a phenomenon is experimentally demonstrable when he knows how to design an experiment so that it will rarely fail to give a significant result.”

So, again, this is as much about experimental design as about the analysis of any given dataset.

Fisher also wrote, “He should only claim that a phenomenon is experimentally demonstrable when he knows how to design an experiment so that it will rarely fail to give a significant result.” Unfortunately, current statistical practice with all its p-hacking allows researchers to reliably attain statistical significance even in the presence of pure noise. Just ask Daryl Bem and Brian Wansink.

In any case, Robinson and Wainer continue:

Fisher believed NHST only made sense in the context of a continuing series of experiments that were aimed at nailing down the effects of specific treatments. . . . NHST as it is used today hardly resembles Fisherís original idea. Its critics worry that researchers too commonly interpret results where p > 0.05 as indicating no effect and rarely replicate results where p < 0.05 in a series of experiments designed to confirm the direction of the effect and better estimate its size. This conception is of a science built largely of single-shot studies where researchers choose to reach conclusions based on these obviously arbitrary criteria.

Yup.

They follow with a bunch of reasonable statements about significance testing and null hypothesis testing that I won’t repeat here because we’ve been thinking about them a lot in the post-2010 world.

There’s one bit at the end of Robinson and Wainer’s article I want to highlight, though, if only because I think it contains some insight and some confusion. They write:

Recently, Jones and Tukey (2000), expanding on an old idea (e.g., Lehmann, 1959; Wald, 1947), suggested a better way in which one could interpret significant and nonsignificant p values. If p is less than 0.05, researchers can conclude that the direction of a difference was determined (i.e., either the mean of group one is greater than the mean of group two or vice versa). If p is greater than 0.05, the conclusion is simply that the sign of the difference is not yet determined.

I agree that this is something that people do, and indeed Francis Tuerlinckx and I made use of that perspective in our 2000 article, “Type S error rates for classical and Bayesian single and multiple comparison procedures,” which began:

The part where I disagree with Robinson and Wainer is where there they recommend interpreting p-values in this way. To say that, just because the p-value is less than 0.05, you can “conclude that the direction of a difference was determined” . . . . I think that’s a mistake. We’ve just seen too many examples where this is not appropriate, starting with that beauty-and-sex-ratio example and going on from there over the years.

Summary

I like the Robinson and Wainer article. It reminds me that, back around 2000, there was a lot of discussion in psychometrics regarding the problems with null hypothesis significance testing. I remember Dave Krantz showing me his article from around that time on the null hypothesis testing controversy in psychology (see brief discussion here) and me thinking that, yeah, this is a problem. But it was only when it combined with the forking-path issue (as noted by psychologists Vul, Pashler, Francis, Simmons, Nelson, Simonsohn, etc.) around 2010 that the true scale of the problem became clear. The Robinson and Wainer article gives a clear perspective on many issues that we’re still discussing today.

23 thoughts on ““On the Past and Future of Null Hypothesis Significance Testing”

    • Sorry, correcting a typo above:

      Checking my understanding: you’re happy with “p > .05 ⇒ direction not determined” but not “p < .05 ⇒ direction is determined”?

      (This does seem reasonable.)

  1. Not to dismiss discussions in 2000 in psychometrics, but discussion of the null in psychology by methodologists happened long, long before then. As just a few examples:

    Rozeboom, W. W. (1960). The fallacy of the null hypothesis significance
    test. Psychological Bulletin, 416–428.

    Lykken, D. T. (1968). Statistical significance in psychological research. Psychological
    Bulletin, 70, 151–159.

    Meehl, P. E. (1978). Theoretical risks and tabular asterisks: Sir Karl,
    Sir Ronald, and the slow progress of soft psychology. Journal of
    Consulting and Clinical Psychology, 46, 806–834

  2. It’s curious that Robinson and Wainer say that the difference between Fisher’s ideas and those of Neyman and Pearson “are sufficiently fine that modern users lose little if they ignore them”, but then go on to detail ways in which their ideas importantly differ! These important differences have been explored by many, notably in my mind by Gerd Gigerenzer at al. in The Empire of Chance (Cambridge, 1989; especially chapter 4, “The inference experts”) and by Richard Royall in Statistical Evidence: A Likelihood Paradigm (Chapman & Hall, 1997). Fisher derided Neyman and Pearson’s conception of hypothesis testing as an “acceptance procedure”, suitable for large enterprises to make decisions on whether to accept a supplier’s products, but wholly unsuited for scientific investigation; he had no use at all for “types” of errors. Statistics textbooks for biologists (and, in my limited experience statistics texts for others) invariably meld the ideas of Fisher and Neyman & Pearson into a gemish that obscures these salient differences, presenting them as “the” method of “hypothesis testing”.

  3. “Fisher understood science as a continuous process and viewed NHST in that context. He often used NHST to test the potential usefulness of agricultural innovations.” It seems worth mentioning that Fisher worked for a long time at an agricultural experiment station where a lot of experiments were going on simultaneously on small plots, and where repeating experiments was easy.

    • | and where repeating experiments was easy

      Depending on what you mean by separate ‘experiments’ and ‘repeating’, I’m not sure I understand.

      If the idea here is that in a given year at Rothamsted you might have 100s of small plots with different treatments. Sure you can test a lot of different treatments simultaneously, and even have replications of the same treatments in that data. But if there is something that unifies these treatments (same crop) wouldn’t you analyze all that data together and not treat the replication as repetition?

      So at the end of this growing season you get a chance to see what treatments are ‘worth’ repeating *next year*. Not only do you have to wait a year to be able to repeat the experiment. After ‘repeating’ the experiment the second year, you realize there may actually be an interaction of the effect of interest with weather and realize you may actually need to run these treatments for ten more years to get a result that generalizes across weather scenarios, but perhaps doesn’t generalize beyond Rothamsted soils.

      Nothing seems to make ag experiments particularly easy to repeat.

      • It does come down to what you mean by repeat. But in the sense that you might do hundreds of plots, with 10s of treatments and 10s of replicates, and then maybe be able to redo the experiments 2-3 times in a growing season that’s a lot more replication than you could do on say nuclear powerplant meltdown containment devices.

        • I’m not sure that this level of replication is that realistic. Often agricultural in-field experiments use relatively large plots (both for required management using large machines and to minimise edge effects), and the crop species have specific growing seasons (winter wheat, spring barley etc.). So, for generalisable field trials, it’s a lot harder to replicate that you might assume (of course, growth room work might be entirely different, but that wasn’t typically the type of thing Fisher was involved with as far as I know). Fisher might also have been involved with distributed experiments across volunteered farms though, Rothamsted did that type of thing too.

        • Yeah, and I’m pretty sure the context of ‘repeat’ throughout these Fisher quotes is deciding what to do next after obtaining yield results and calculating the p-value.

          So the number of plots is only relevant to how many different treatment effects you can simultaneously estimate and the precision of the estimates and not to the degree to which that experiment can be repeated.

          I don’t know anything about growing potatoes in England… if you tell me you can get three harvests from a plot in a year at the time of Fisher I’d believe you, but I’d be skeptical that these three growing seasons would a priori even be considered climatologically similar enough that you can have a repeat in the same year.

          And if your goal is effect estimates that generalize across weather scenarios, your “experiment” cannot consist of just a single year. And maybe you’ll really only get a chance to consider repeating it in 5-10 years.

  4. I really like this discussion of Fisher. It foregrounds the institutional structure of experimental replication.

    In my fields (social psychology, sociology), I worry we often delude ourselves about that structure. We *want* to believe that researchers execute replications for what we see published. If a study has been around long enough (and has enough citations) or (groan) hits the news, surely *someone* (we fantasize) has subjected the study to the heat of independent verification. People with PhD’s who have published experiment-based research believe this!

    I’d be delighted if a qualitative researcher interviewed experimentalists in different disciplines/areas (industry, STEM, social, biz) just to see what sorts of beliefs about experiments & replication are floating around out there.

  5. What’s great about this is that Fisher’s paper is nearly a hundred years old but a large proportion of the scientific community is still happily ignoring what he said!

    Why is it that no one ignores Einstein but everyone ignores Fisher, even though it’s been shown over and over and over again that statistical significance as expressed by Fisher is just a first pass low accuracy test? I like that Fisher emphasizes experimental design, because that’s the key to whether “significant” results are likely to prove up in the long (or even the short) run: if the experiment is poorly controlled and depends on weak or – more commonly – unrecognized – assumptions, the results are unlikely to prove up. however if the experiment is tightly constrained to a single dimension of variation, the chances it will prove up are much better.

    So after all the statistics actually *is* easy. It’s the science – the part about getting your experiment and formulas to represent the real world – that’s hard. And maybe that’s why statistical teaching fails so badly: it teaches only statistics.

  6. For the Jones and Tukey-style interpretation of tests, where one makes a decision on the sign of the parameter but not more than that, you might like this recent review. Some of the tools to assess testing decisions once they’ve been made may be of particular interest. There is, of course, more work to do.

  7. Andrew today mentions

    “Simmons, Nelson, Simonsohn”

    They, it turns out, are in the thick of uncovering the current Harvard University problem regarding data manipulation. Because of this, I have been reading the book, “Rebel Talent”, by Francesca Gino, the person most under fire. Putting aside any issue of falsification of data, the book is very well written and engrossing. Nevertheless, she seems to be a true believer when it comes to: separate the subjects (somehow, but not necessarily at random) into two groups, treatment and non-treatment, respectively, and

    low p-value implies “significance”

    which is not just applicable to Harvard undergraduates but to the universe as a whole.

    One of her examples (page 16) has to do with wearing red (Converse) sneakers in one section of a Harvard class and more traditional foot gear in the preceding class. Her conclusion from this successful experiment is “…we all share the desire to be happy…And something as simple as a pair of red sneakers might make all the difference.”

Leave a Reply

Your email address will not be published. Required fields are marked *