John Carlin writes:
I wanted to draw your attention to a paper that I’ve just published as a preprint: On the uses and abuses of regression models: a call for reform of statistical practice and teaching (pending publication I hope in a biostat journal). You and I have discussed how to teach regression on a few occasions over the years, but I think with the help of my brilliant colleague Margarita Moreno-Betancur I have finally figured out where the main problems lie – and why a radical rethink is needed. Here is the abstract:
When students and users of statistical methods first learn about regression analysis there is an emphasis on the technical details of models and estimation methods that invariably runs ahead of the purposes for which these models might be used. More broadly, statistics is widely understood to provide a body of techniques for “modelling data”, underpinned by what we describe as the “true model myth”, according to which the task of the statistician/data analyst is to build a model that closely approximates the true data generating process. By way of our own historical examples and a brief review of mainstream clinical research journals, we describe how this perspective leads to a range of problems in the application of regression methods, including misguided “adjustment” for covariates, misinterpretation of regression coefficients and the widespread fitting of regression models without a clear purpose. We then outline an alternative approach to the teaching and application of regression methods, which begins by focussing on clear definition of the substantive research question within one of three distinct types: descriptive, predictive, or causal. The simple univariable regression model may be introduced as a tool for description, while the development and application of multivariable regression models should proceed differently according to the type of question. Regression methods will no doubt remain central to statistical practice as they provide a powerful tool for representing variation in a response or outcome variable as a function of “input” variables, but their conceptualisation and usage should follow from the purpose at hand.
The paper is aimed at the biostat community, but I think the same issues apply very broadly at least across the non-physical sciences.
Interesting. I think this advice is roughly consistent with what Aki, Jennifer, and I say and do in our books Regression and Other Stories and Active Statistics.
More specifically, my take on teaching regression is similar to what Carlin and Moreno say, with the main difference being that I find that students have a lot of difficulty understanding plain old mathematical models. I spend a lot of time teaching the meaning of y = a + bx, how to graph it, etc. I feel that most regression textbooks focus too much on the error term and not enough on the deterministic part of the model. Also, I like what we say on the first page of Regression and Other Stories, about the three tasks of statistics being generalizing from sample to population, generalizing from control to treatment group, and generalizing from observed data to underlying constructs of interest. I think models are necessary for all three of these steps, so I do think that understanding models is important, and I’m not happy with minimalist treatments of regression that describe it as a way of estimating conditional expectations.
The first of these tasks is sampling inference, the second is causal inference, and the third refers to measurement. Statistics books (including my own) spend lots of time on sampling and causal inference, not so much on measurement. But measurement is important! For an example, see here.
If any of you have reactions to Carlin and Moreno’s paper, or if you have reactions to my reactions, please share them in comments, as I’m sure they’d appreciate it.
I find the pedagogical approach useful: emphasize the research question and distinguish between description, prediction, and causation. Beyond that, however, I’m not sure I learned much from the paper. There is a sort of “catalog of forked paths” feeling I get from the various ways in which regression models may be misused. These are useful cautions. But I never quite understood what the alternative was: if the regression model is misspecified or misused, what are the authors proposing instead? Obviously, improve the analysis and don’t misinterpret (hype or otherwise overstate the results of the model), but are they suggesting that some other type of analysis be used? If so, what is that? I did read the whole thing, although perhaps not as carefully as needed if the answers were there and I just missed them. But the title seemed a bit off to me – it appears to suggest where regression models are useful and where they are not, but I didn’t find that in the paper. What I found instead was a useful explanation of the many ways in which things could go wrong.
In my teaching, I try to achieve this by exposing students to rich multivariate data and having them build many models (mostly regression models) to gain an appreciation for the many assumptions that must be made, and the number of forked paths followed. The result, I hope, is a more humble appreciation for any particular model they may produce, and a more informed consumer of any model they are presented with. Other than paying more attention to defining the research question and the nature of the three types of uses (along with Andrew’s three types of generalizations), I am not seeing what I would change on the basis of what the paper says.
First two sections were good, but the third fails to address the main problem with the “causal” purpose:
In the vast majority (I’d guess 99.9+%) of cases:
1) There is no basis for choosing one functional form over another. So it is arbitrary, thus so are the coefficients.
2) Data is not available to do this “full adjustment”.
Therefore this is not a valid use-case for regression models. Instead they generate conflicting and misleading information that wastes the careers of generations of researchers.
The solution is to actually work out a theory (set of premises) of what is going on, then derive your quantitative model from those assumptions.
I think we’ve been here before, but I must be particularly dense to keep not getting it.
Use case: I want to understand what contributes (causative statement) to unemployment of a particular demographic group. I’d really like to do some experiments, but they would be very costly and take a long time and nobody is willing to give me the required money for these. But I do have considerable observational data: time series and cross sectional. I know that the macro economy changes over time and is relevant. I know that the nature of employers (industry, firm size, etc.) varies across geography, as do regulations (such as minimum wages and other employment laws and regulations). I have some idea of the qualifications (educational background, years of work experience) of workers in different places and over time.
I can and should be more specific about these things, but I don’t think my reasoning would qualify as a data generating process. I’m not sure I see how to “derive your quantitative model from those assumptions” beyond identifying relevant variables. Are you saying this is not a valid use case for regression models? If not, what would you propose? And, if the functional form needs to be derived from a theory (“set of premises”), what theory would provide this functional form? Somehow, I feel that I’d rather let the data guide the choice of functional form than trying to specify it from whatever theory I might invent.
What Anoneuoid is saying is that if you have a theoretical reason to select a particular regression blender, and you make surprising predictions (sans the model) that pan out, then we might really have something there! We can say, without lying to ourselves, that those regression coefficients probably contain something of interest. But simply grabbing coefficients from a model of convenience and interpreting them as meaningful is a delusion.
Knowing the relevant variables is not enough. Coefficient values from a regression are conditional on the model specification and there are infinite competing specifications once could select from that would produce conflicting values.
I don’t disagree that choosing a random set of predictors and interpreting the coefficients as meaningful is bad. But choosing variables that make sense from a theoretical view (for example, the factors that would influence income levels), and avoiding known serious omitted variable issues, is often the best we can do. If those coefficient values are meaningless, then it sounds like you are saying that there is no useful analysis for analyzing factors contributing to wage rates for demographic groups (such as trying to see if there is evidence of discrimination). Once we rule out a controlled experiment, I’m not sure what else can be done in this case. Of course there are an infinite number of competing specifications – I would argue there always are (at least in social science research). What Anon has argued in the past – and what I still don’t accept – is that because there are an infinite number of possible specifications it means that no one specification can be interpreted.
I understand the quote differently and I think your reaction to it is in fact consistent with the point provided in the quote.
Regression model itself never provides “the truth” (eg. value of estimated regression coefficient =/= true average causal effect), unless it is conditioned on some “true causal structure”. The problem is we (almost) never know the “true causal structure”, therefore you derive one somehow; be it from a theory, your own intuition or w/e, and then you act as-if it is a true causal structure and interpret the coefficient according to this assumption (eg. this b1 = average causal effect | my causal structure is true). The interpretation of a single coefficient is always context dependent.
Now, the problem – if the causal structure you created assumes a certain functional form of used variables, interactions between them etc., but your data doesn’t provided enough information, you can’t identify and estimate the coefficient of interest.
As a journal editor, the biggest problems I see have to do with over-use of regression models. A general problem is failing to distinguish between a measure (variable, predictor, …) and what it is trying to measure. This distinction is important when the question concerns mechanisms of effects, e.g., in accounting for individual differences in some cognitive bias. The problem goes beyond questions about reliability of predictors, because the “true score” of a test (in the sense of classical test theory) is not always the same as the underlying construct that the test seeks to measure. Researchers don’t seem to realize that regression coefficients depend on: the reliability of predictors; which predictors are included in the model; the coding of variables when interactions are included in the model; and what we might call the true validity of measures (tests), that is, whether a test measures something other than what it is trying to measure. (The most blatant examples of the last are the use of “proxy” measures, such as zip code for wealth. But then “wealth” itself is often a proxy for something else.) They also don’t understand that interactions are sometimes “removable” by transformations of a scale of measurement (e.g., replacing probability with log odds).
I see these problems in other journals, sometimes those with huge “impact factors” like the Journal of Personality and Social Psychology. Often the abstract alone is sufficient, when it says something like “X predicts Y, even when controlling for Z.” Only very rarely does the paper consider the error of measurement in Z.
I did not read the paper you cite. But I think I agree with some of what people say. I also agree that regression models are sometimes completely appropriate when used in the standard way, e.g., developing a formula to predict various health events from measures easily available.
Someone who was trained in integer optimization before math stats, I think we should teach both before one embarks on regression. Regression is on the real plane, where’s integer optimization is used to find shortest path type network problems. Regression depends on convex set assumptions, where’s network optimization is non-convex by definition (otherwise we will solve shortest paths with straight line geometry). The forking path is essentially a non-convex, network problem and appears to be a thorn in statistics. Actually, non linearity, complexity, forking paths are all part of complex systems (non-convex sets). Regression focuses on linearity/linearized models as convex only solutions. Bayesian networks, with priors and hyper-priors, allows a network approach and can be used in hierarchical data (eg geographic data). The adjacency matrix provides for a (non-convex) conditional/serial autoregressive, linearization approach.
Very nice paper. It reminds me somewhat of Jim Hodges’ paper on statistical practice as argumentation: https://www.biostat.umn.edu/~hodges/HardToFind/Argumentation_1996.pdf
This article reminds me of similar points raised by David A. Freedman in “From association to causation via regression”: https://statistics.berkeley.edu/sites/default/files/tech-reports/408.pdf and “Statistical models and shoe leather” (https://psychology.okstate.edu/faculty/jgrice/psyc5314/Freedman_1991A.pdf). In the latter, he discusses John Snow’s careful multifaceted approach compared to someone else, who used the modern (I’d say also lazy) standard regression approach. My view (I work in epidemiology) is that any scientific investigation should go beyond the data (whether observed or simulated) and the choice of statistical models. Inherently, because of their profession, this is what statisticians end up discussing, but for me, an investigation should look at a question from several angles, and should be based on more than numbers. But yes, an intresting paper, food for thoughts.
The Freedman article is excellent. It has done more to persuade me than the paper in Andrew’s post above or the arguments raised by Anon, AllanC, and others. Of the many quotable statement in Freedman, is this gem:
“Regression models often seem to be used to compensate for problems in measurement, data collection,and study design.”
I can’t argue with that. I’m not sure I am willing to go as far as Freedman is in rejecting regression models to study causation – not because of any technical improvements, but because I think he underestimates the difficulties of using better methods in many areas of social science (he does, however, address this point). Much of economics deals with issues where experimentation is difficult and costly – indeed, experiments have their own issues (generalizability). Observational data is often available and in large quantities – I find it hard to conclude that we must ignore its use for studying causation. At the same time, I totally agree with the statement above – it cannot and should not substitute for problems in measurement, data collection and study design.
One thing Freedman does not address is validation. Regression models are rarely subject to validation – it is rare to find a regression model where the data is split into training, validation, and test data sets while that is common practice in data science. I recall some discussion of this on this blog (though I can’t seem to recall where). I haven’t seen any good reason why validation of regression models is seldom used – it seems like a case where the technical regression model assumptions are taken to substitute for actual testing of the model’s implications.
I’m back to near where Freedman starts: “Questioning the value of regression is then tantamount to denying the value of data.” He roundly rejects that argument, but that is where I am left. I’m convinced by his critique, but unsure where that leaves me. So much research uses regression modeling in some form, and so much observational data is available relative to good experimental data, that it isn’t clear to me what to conclude. If we are to reject most causal reasoning that uses regression models, then where does that leave the myriad policy decisions that must be made? Are we suggesting that these decisions should be made based on intuition until better experimental designs and measurements are developed that can more reliably address causation? Complicating even that “solution” is the fact that such designs are no panacea for so many questions that need answers – do social media cause depression, does gun control reduce gun violence, do minimum wages increase unemployment, does exercise of type X reduce medical condition Y, …?????
Linear models absolutely *can* be used to study causal questions, including with observational data, you just need to acknowledge that you have to pose a causal inquiry via your model/assumptions…it’s not a case where “let the data speak” really works. And it has nothing to do with the kind of test/train split that is used to check *predictive* quality.
I highly recommend Richard McElreath’s treatment of causal inference. It differs slightly from Andrew’s in that he is a big fan of DAGs, but as a pedagogical tool for us visual learners there is nothing better IMO.
Also I recommend chapters 18-21 of Regression and Other Stories.
Andrew
Those chapters in your book are quite helpful. I am still confused about 2 different, but related, issues that have come up in these comments and readings. One concerns using observational studies to make causal inferences. There are quite a number of difficulties, none insurmountable, but all requiring much care. I think everyone agrees that it is possible to do so, although opinions will vary concerning just how effective various tactics (e.g., DiD, instrumental variables, missing variable bias adjustments, etc.) are. As you point out, even data from randomly designed experiments have some of the same issues – if the confounding variables are not also randomized.
The other issue is the functional form of the regression models. A number of people point to the need to derive this from theoretical models, a worthy goal in my mind, but essential according to others. Yet, even in your careful analysis, I don’t see the functional form of the regression models as being derived from any first principles. Interpreting the coefficient in a regression model as indicating a causal relationship is hazardous, doubly so (at least) in observational studies. But even in a perfectly randomized experiment with randomized confounders, there is still an issue of whether the regression model should be linear, what interactions should be modeled, exactly which variables need to be included, etc. These are all real issues, but I am having a hard time seeing how these choices need to result from a theoretical model rather than relying on an inductive approach from the data.
For many social science questions, I don’t see requiring the functional form to be derived from a theoretical model as superior to fitting various models to the data and comparing their fits to the data.
To make this concrete: a demand model could be constructed for a good, say an electric vehicle. One approach would be to specify a utility function with a specific functional form, based on various characteristics (price, speed to 60mph, horsepower, size, etc.). A demand function can be derived from this specification. Alternatively, those same characteristics could be used to fit demand (assuming we have a data source providing data on consumers who have bought an EV and their various characteristics, as well as the car’s characteristics) – trying various functional forms for the demand regression model and comparing these. I’m not sure I see the model-derived demand function as a better approach – I have no idea what an appropriate utility function looks like, beyond my ability to list the variables that it would depend on. So, I’m having trouble seeing how the call for deriving regression models from theories relates to the issues of whether and how to interpret regression coefficients causally.
I have certainly experienced use cases where properly specifying the model is important (and where people i.e. reviewers didn’t care whether I properly specified the model). On the other hand, to take Dale’s electric vehicle example to the extreme, the rapid growth of use of deep neural nets illustrates that there are plenty of use cases where not only do you not worry about specifying the model, you don’t even get to know what the model is.
“We then outline an alternative approach to the teaching and application of regression methods, which begins by focussing on clear definition of the substantive research question within one of three distinct types”
This idea is central in Richard Berk’s 2004 Sage book, Regression Analysis: A Constructive Critique, which, as you might expect from the title, has much to say about problems in the way models have been used. Berk’s book has influenced my thinking and teaching, and I share some sections from it with students every year. The first couple of chapters are great for getting a handle on the essence of regression, and the big message throughout the book is that the analytic goals matter. While I don’t think that these ideas are new for teaching in 2024, I think it would be great if more folks were exposed to them, as problems persist :).
David:
Yes, I reviewed Berk’s book a few years ago.
Thanks Andrew for posting a pointer to our preprint. It’s great to see some discussion here, but I feel some of the key points have been missed. This is no doubt in part due to a few gaps in the paper, which we are trying to fix in a revision currently underway (which should be available in the next few weeks).
Several reviewers (and commenters) have implicitly highlighted the fact that we were not sufficiently clear that although the paper starts from the fact that regression models are ubiquitous and are widely misused and misunderstood, our proposal is NOT to develop a new approach to teaching regression per se. The idea that the method or technique is central is the fundamental problem that we identify, and we propose that this can only be addressed by turning the initial focus of teaching to the purpose or research question. This needs to be done very seriously, i.e. there should be no “general theory” of regression models, or of how to fit them “well”. There will always be some applicable mathematics behind the application of regression methods, but by starting with the math, traditional courses immediately deflect the students’ attention away from the purpose, leading to all the problems that we identify under the heading of the “true model myth”.
There has been a longstanding tendency to identify the role of the biostatistician with models and estimation techniques, undervaluing the essential role of statistical thinking in framing well-specified research questions. In contrast, our key assertion is that the three types of question need to be taken seriously, at the beginning of every analytic investigation. Identifying the type of question requires considerable reflection and discussion between statistician and collaborator. Once identified though, very briefly, if you have a prediction question, you might then think about tools A and B for developing prediction algorithms (multivariable regression being a great one if done well)… If you have a descriptive question, the first thoughts for analysis might be simple descriptive statistics, but there might be complicating issues to attend to (re sampling bias, perhaps, for which some regression technology might help). If you have a causal question, then you need to get serious about defining the question precisely, in terms of a target parameter or estimand, before thinking about the potential role for models to assist in estimating that quantity. Rather rarely, we suggest, will this target parameter be naturally defined by a coefficient in a regression model, although under carefully considered assumptions a regression coefficient might be relevant. Finally, taking the question seriously implies serious questioning as to whether aims such as “identifying risk factors” are meaningful.
I would urge careful reading of the paper (long as it is, sadly, and the next version will be a bit clearer) before concluding that it’s all been said before, although of course others have clearly made related observations.
Andrew, “Regression and Other Stories” gets a citation in our Discussion, with the observation that it “remains ambiguous about where regression models come from – in particular, whether a model specification may precede a purpose.” Perhaps this was too harsh?!
Finally, just to note that my colleague uses her full name “Moreno-Betancur”, to distinguish from all the other Morenos out there!
Thanks, John and Andrew, this is a very important discussion !
I definitely prefer research and teaching that
“begins by focussing on clear definition of the substantive research question”
versus the less-focused
“build a model that closely approximates the true data generating process”
Regression and Other Stories is great. It does have a few places where more purpose clarity would really help. For example, in Chapter 11 about model evaluation, section 11.5 uses predictive simulation to check the fit of a first-order autoregression to yearly unemployment data. Visual and numerical checks show lack of fit, but what is the research question ? do we care about this lack of fit ? or is the model *good enough* to answer our research question ?
I have been reading the functional estimation literature (see this fun book: https://alejandroschuler.github.io/mci/, Kennedy’s excellent review: https://arxiv.org/abs/2203.06469), where they target analysis to a specific estimand. Their asymptotics-based approach may not be practical in my context where I need to include prior information to improve estimation in small groups. But I really appreciate the clarity of a target estimand.
Thanks Shira! You make a great point about the gap between the asymptotics-based focus of a lot of modern causal inference and the real-world problems of subgroups and sparse data. I think we need a lot more work on bridging that gap — but indeed, keeping the estimand(s) firmly in view before elaborating the models.
Oops, sorry forgot to self-identify! That anonymous was me…
It seems odd to me to view description as one of the main tasks that regression models may serve. Why not simply rely on Pearl’s Causal Hierarchy that delineates, in a well-defined and deeply theoretically justified way, the three most important, fundamentally distinct classes of tasks that quantitative models – including regression models – may serve: predictive, interventional, and counterfactual/structural?
For instance, a linear regression model
– can *always*, albeit with varying success, be used to predict stuff, even if it is descriptively problematic (“very false”)
– it may only *sometimes* – under well-defined conditions – be used to estimate interventional quantities
– it is *almost never* a good candidate structural model, which can be seen immediately by imagining that the true values of its coefficients are given and by considering the typically absurd counterfactual consequences of viewing this (kind of very simple) model as a structural model.
Apologies if this wasn’t clear but we do not regard “description as one of the main tasks that regression models may serve”, but description is one of the three types of research question that an empirical investigator may seek to answer. Regression may or may not be helpful with aspects of that task, e.g. to do some kind of regression standardisation to a different population distribution from the one that we sampled from… Once again, we are starting from the tasks and asking which methods may contribute (to each), not starting from the methods and trying to figure out which tasks they may serve. I do generally agree with your summary of what a regression model can do — it’s just that the structural “task” is not an actual empirical question in our intended meaning of “questions” or “data science tasks”; I would view structural models as tools within causal inference.
Thank you for the clarification! So your main objection has to do with the status of counterfactual/structural questions. It is interesting that you do not view the “structural task” as an empirical question or a “data science task,” but you view interventional questions as empirical. I imagine you agree that to answer an interventional query, one must typically introduce:
– testable statistical assumptions
– testable causal assumptions
– untestable statistical assumptions
– untestable causal assumptions
When the interventional conclusion is data-dependent, the question/task is empirical, since empirical simply means observation- or data-dependent (it sure does not mean theory-independent).
What about the empirical character, or lack thereof, of counterfactual/structural questions, then? It may seem that counterfactual questions are not empirical by definition since they specify conditions that cannot ever happen. However, even though the true structural model *for a particular set of modeled (i.e., endogenous) variables* is underdetermined by observational and interventional data *in these variables* (as follows from the PCH theorem) which makes the task of identification of such a model in some special sense “nonempirical”, a structural model (that provides answers to all counterfactual questions) may be sufficiently constrained by *other kinds of data*, i.e., when *additional variables* are observed or intervened on. Otherwise, it would not make sense to claim that e.g., a particular structural model, such as e.g. a structural model of a circuit, is approximately true, right?
The three distinct uses identified are broadly the same as those identified by Tredennick et al. 2021 in their guide to model selection for ecologists. Their paper is probably the best I’ve seen on the topic and closely overlaps with this one. Something Tredennick et al. do that this paper might benefit from is clearly defining when the 3 uses can and cannot be combined.
Otherwise, great paper! The focus on teaching is great. Agreed with Gelman that even in PhD level statistics courses, getting students to build intuition around the basic deterministic piece of regression, i.e., y=b0+b1*x is a major hurdle. In my experience, once students understand this bit, building intuition around error comes easy.
I’ve consistently found (as widely suggested by others) that having students simulate their data, then use regression to recover the true values (or not!) is extremely useful for building their understanding of when and how inference can go astray. Perhaps the paper could incorporate the practical suggestion of simulating data as a core part of teaching.
Tredennick, A. T., Hooker, G., Ellner, S. P., & Adler, P. B. (2021). A practical guide to selecting models for exploration, inference, and prediction in ecology. Ecology, 102(6), e03336. https://doi.org/10.1002/ecy.3336
The three uses that Carlin proposes are fairly frequent in the literature, and existed long before Tredennick et al. For example:
Hodges (1996) https://www.biostat.umn.edu/~hodges/HardToFind/Argumentation_1996.pdf
Hernan et al. (2019) https://www.hsph.harvard.edu/wp-content/uploads/sites/1268/2019/04/hernan_chance19.pdf
Shmueli (2010) https://www.stat.berkeley.edu/~aldous/157/Papers/shmueli.pdf
Gelman et al. in Regression and Other Stories also have a similar three-part classification and I’m sure there are many others.
I don’t personally think that the categories of Tredennick et al (“data exploration, inference, and prediction”) are very clear. “Inference” is vague as they define it, (“to evaluate the strength of evidence in a data set for some statement about nature”) and conflates a number of specific tasks: for example, inferring the presence of an effect via something like Neyman-Pearson hypothesis testing is decision-theoretic, but inferring likely values for a regression coefficient, without referring to P-value cut-offs, could be either descriptive or causal, depending on the task at hand. It doesn’t make sense to me to say that estimating some beta in a causal problem is inference, but predicting some values of Y is not inferring something about the real world. Labelling these things as descriptive, causal and predictive inference is far more intuitive to me in terms of clear communication. I like the Hodges (1996) classification that clearly separates out decision-theoretic elements of statistical arguments from their inferential parents.
“including various “non-linear” terms such as polynomial functions to represent curved relationships”
From reading ‘Statistics for Experimenters’ I always understood that these surfaces were essentially Taylor Series expansions of the assumed surface which is why we see x^2 and x^3 terms rather than treating the exponent as a parameter itself which could be 1.8 for example.
I’ve always had some difficulty in imagining how one could interpret polynomial functions causally apart from to say there is some form of curvature in the relationship.