“Why did ANOVA fall out of fashion?”

A student asks the above question.

My response: Anova is still important; it’s just been subsumed by hierarchical models.

The link is to my 2005 paper, Analysis of variance: Why it is more important than ever, which begins:

Analysis of variance (ANOVA) is an extremely important method in exploratory and confirmatory data analysis. Unfortunately, in complex problems (e.g., split-plot designs), it is not always easy to set up an appropriate ANOVA. We propose a hierarchical analysis that automatically gives the correct ANOVA comparisons even in complex scenarios. The inferences for all means and variances are performed under a model with a separate batch of effects for each row of the ANOVA table.

We connect to classical ANOVA by working with finite-sample variance components: fixed and random effects models are characterized by inferences about existing levels of a factor and new levels, respectively. We also introduce a new graphical display showing inferences about the standard deviations of each batch of effects.

We illustrate with two examples from our applied data analysis, first illustrating the usefulness of our hierarchical computations and displays, and second showing how the ideas of ANOVA are helpful in understanding a previously fit hierarchical model.

This is the paper that discusses the five different definitions of fixed and random effects from the literature; scroll down to page 20.

I guess the problem is that Anova got associated with clever-but-ultimately-bad ideas of null hypotheses and F tests. Hierarchical modeling remains very important; I think it’s an underrated topic in statistics.

17 thoughts on ““Why did ANOVA fall out of fashion?””

  1. typo:
    “Anova got associated with clever-but-ultimately-bad ideas of null hypotheses and F texts.”

    Should be

    “Anova got associated with clever-but-ultimately-bad ideas of null hypotheses and F tests.”

  2. What even really qualifies as ANOVA? When people look at the word ANOVA what are they thinking about or referring to?

    (note, it’s not that I don’t know “what ANOVA is” it’s more that to me ANOVA is used to refer to several things that I think of as distinct and I’m not sure which things people most commonly mean)

    Here is what Wikipedia says: “If the between-group variation is substantially larger than the within-group variation, it suggests that the group means are likely different. This comparison is done using an F-test.” so basically, that article argues it’s a way to test “if the group means are different” and I find that absolutely useless but also I think sometimes people say “ANOVA” when they really mean “linear regression with 0/1 predictors” and they are estimating the coefficients (group means). Also there’s kind of orthogonal basis function decompositions, in this sense Fourier analysis is basically ANOVA… so what do people think?

    • In most cases I’ve seen, ANOVA is interpreted as a *test* of equality of means. My non-statistician colleagues rarely think of estimating means when applying an ANOVA. That estimation is reserved for “post-hoc” analyses, and only when “ANOVA is significant.”

      • Thanks, glad to get some feedback. ANOVA in this context is just harmful. In essentially every situation where you would run ANOVA as a test you should instead just run a linear model and get coefficients and confidence intervals. Also you should include into your model any continuous predictors that might help account for known causes of differences.

        • Daniel, they are all the same thing, GLM, general linear models. So by doing ANOVA you are running a linear model and getting coefficients and CIs. There is nothing ‘harmful’ in ANOVA. Quite the opposite. Wherever you see it used, it will be preceded by a solid experimental design with highly controlled variables. Regression which is used for predictions where conditions were not or could not be manipulated/controlled. Predictors are categories of nominal variables in ANOVA. Horses for courses.

        • In ANOVA table output ive seen, it simply doesnt give coefficients or CIs,it gives sums of squares and F test p values. so while the math is the same underlying it, the information given is harmful because it focuses on a question. “are the means of the groups statistically significantly different” and thats just never the question you want to ask.

          Are the means of the groups different always always always has the answer “yes” (its not even a little bit credible that two groups could have the same mean for the first 1000 digits, much less infinite digits). So the question is always about can you detect the difference? and thats entirely a question of sample size and effect size,but traditional ANOVA output elides effect size estimates entirely. this is the harmful aspect. it teaches people to ask the wrong questions because theyre framed in terms of tests of hypotheses

    • I agree that the term ANOVA is used differently in different contexts. I’m not quite clear what going out of fashion means, either. I think “out of fashion for what?” is a relevent question. Are we talking about R^2 not being a (central) feature of published analysis that is extensively discussed? Or are we talking about helpng students understand the idea of signal and noise in a model? Agree that the Wikipedia description seems really unhelpful.

  3. Why would ANOVA lose popularity if it is mathematically equivalent to regression analysis? The same people who would rely only on p values in one would do the same in another, if that’s the issue. Just as r2 is a form of effect size, so is eta squared and partial eta squared.

    ANOVA became popular because of experimental design where researchers tightly control the conditions, while regression is used in ‘found’ data, post hoc observational studies, where no control is possible. Economics is one example. The beauty of ANOVA is in control of variables and research design that preceded the analysis.

    I am not sure how exactly to set up a 2x2x3 regression for a very well-organized and controlled study, so the output and explanation makes sense to anybody. It is called ANOVA because while it compares means it ANALYSES variances (b/w and within groups/levels).

    I am open to being very wrong if someone could explain how exactly regression is superior to ANOVA if they are the same under the hood. y=b0+b1x+e. The same people think that logistic regression is used for classification only, not realizing there is a continuum and conditional probability it is based only, not only 0 and 1.

    • Navigator:

      People do say that Anova is mathematically equivalent to regression analysis, but I’d prefer to think of it as an add-on to regression analysis, in which the predictors are structured and we estimate the variances of batches of coefficients.

    • Why would ANOVA lose popularity if it is mathematically equivalent to regression analysis? Well, precisely BECAUSE it is mathematically equivalent to regression analysis. I find it much easier to teach regression analysis. It is, in a sense, much more intuitive. ANOVA looks like some complicated way to get exactly the same results.

      Which also means ANOVA did not lose popularity at all. The word ANOVA lost popularity. But the method still exists, hidden as a special case of regression analysis.

    • Navigator,

      Might partly be an interface thing. In R, the summary function on an lm object gives the regression weights, whereas summary on an aov object gives sum-of-squares estimates. You can get regression weights from the aov object and you can get sum-of-squares from the lm object with an extra step. In my experience in the psychology world, most hypotheses are about linear relationships and mean differences rather than variance. In a simple design, a regression weight gives you the mean difference or the linear relationship whereas a measure of variance accounted for doesn’t.

      I also don’t often see explicit hypotheses about differences among a cluster of groups (e.g. are there substantial differences between the 3 levels of this factor? did this manipulation do anything?). The ANOVA F-tests that I see are often performed out of routine, where the F-test serves as a gatekeeper determining whether the actual hypothesis can be tested with a followup t-test.

      I think the complexity of the design is a separate matter. If you care about means in a complex design, an ANOVA doesn’t really solve anything. My advice is usually to calculate estimated marginal means and predictions from the model. I also tend to discourage ANOVAs in these cases too because many researchers aren’t well trained in using appropriate contrasts or being careful with covariates or being clear with what’s being tested with the F-test when repeated measures across trials is part of the model.

  4. ANOVA was the acronym for ANalysis Of VAriance. In my first graduate school class for statistical methods in the 70’s, we hand calculated the sum of squares for the various components (total, treatment, blocks…). If the design was balanced, it was straightforward, although tedious. I do remember thinking it was impressive that all these calculations actually produced something useful and that the total sum of squares could be broken down. As computers and software developed for unbalanced designs, I learned that ANOVA should be replaced by GLM, general linear model. That was also the sequence of development in SAS, PROC ANOVA in the 70’s followed by PROC GLM in the 80’s.

    I’m not sure ‘out of fashion’ is the right terminology. As experimental design moved away from balance requirements, ANOVA became dated. In today’s world, GLM is also dated because most analyses today recognize that there are several random components, not just random error. Once you consider hierarchical random components, then it just seems natural that the analyses should shift away from frequentist to Bayesian inference.

Leave a Reply

Your email address will not be published. Required fields are marked *