Combining forecasts: Evidence on the relative accuracy of the simple average and Bayesian model averaging for predicting social science problems

Andreas Graefe sends along this paper (with Helmut Kuchenhoff, Veronika Stierle, and Bernhard Riedl) and writes:

We summarize prior evidence from the field of economic forecasting and find that the simple average was more accurate than Bayesian model averaging in three of four studies; on average, the error of BMA was 6% higher than the error of the simple average. We also add new evidence by reanalyzing the results from a study on election forecasting, which was published in Political Analysis and promoted the use of BMA for combining forecasts. However, the authors did not compare the results to the simple average as the widely established benchmark. We did: the error of BMA was 31% higher than the simple average.

I love this sort of empirical evaluation. Here are my quick thoughts:

One way this sort of “overfitting” finding can be interpreted is that you have prior information that is not being used in the model. If a simple average performs relatively well, another way of interpreting this is that you are fitting a model whose coefficients have very very strong prior information constraining them to be equal. My guess is that, in general, a slightly weaker prior would perform even better. But if the prior is too weak, you’re throwing away information and then you get the poor performance identified with BMA above. (BMA is only as good as the models it is averaging.)

11 thoughts on “Combining forecasts: Evidence on the relative accuracy of the simple average and Bayesian model averaging for predicting social science problems”

        • It seems that model combination implies a mixture model, where you could view each observation as being drawn from a different model in the mixture. Whereas model averaging implies the whole dataset is drawn from one of the models, just that we don’t know which. Would that be correct? If so perhaps it gives some intuition about the pros and cons.

    • EBMA (as used in this paper) sounds like one particular incarnation of all the possible ways the component models of an ensemble might be weighted. (I might be wrong)

      Do the competition models use EBMA or something more sophisticated?

      • It’s usually linear regression or something similar. Ie. first use N models to make N predictions and then usese predictions as a input for linear regression (the sum of weights doesn’t have to be 1 in this case, but otherwise almost like weighted average). Something like that was used by Netflix winners (paper: https://www2.research.att.com/~volinsky/netflix/Bellkor2008.pdf ) and in many Kaggle competitions. In general this kind of approaches are called stacking in machine learning literature.

        Idea behind BMC is pretty much the same (link above): first make predictions, then take some weighted combinations of those(=create new models) and finally apply BMA over those new models.

  1. I’m trying to understand Andrew’s comment about the “very very strong prior” constraining coefficients to be equal and overfitting as a culprit for BMA’s poor performance.

    If we have E(Y) = XB, with B in R^p, this prior is over B? So, say, B ~ Normal(b0, sd=.01)? I can see how this would reduce the variance of the coefficient estimates, but how does this relate to using a simple, unweighted average of models, which I interpret as implying a flat prior over the space of models?

  2. I found this snippet from the paper highly non-intuitive:

    The supposition that a method’s past performance is an indicator of future performance may be plausible and appeal to common sense. However,…it is often contradicted by empirical evidence. In many real-world forecasting problems, a model’s past performance is unrelated to the model’s predictive performance in the future.

    If a model’s past performance is *unrelated* to future performance what’s the point behind any validation? Or even behind using any empirical data at all to train or tune a model?

    Might as well build fully open loop models, guided only by fundamentals and theory?

    What gives? What am I missing in my naive understanding?

    • I agree, I think the statement that past performance is unrelated to predictive performance, if interpreted as written, suggests that the whole enterprise (at least in terms of weighting models relative to each other) is a waste of time. I think this statement is probably a bit strong.

      To me a big part of the issue comes down to the relationship between the quantities you are using to score models and the quantities you are trying to predict. If these quantities are basically the same (limiting case would be another sample from the exact same population as used to train the model(s), looking at the exact same outcome), then it might be reasonable to penalize quite harshly those models with poorer fit to the data (overfitting issues aside). On the other hand if the outcome we are interested in is very different to the outcome we are using to evaluate the competing models, then it is probably dangerous to penalize them too harshly.

      This was my takeaway from arguments by Andrew and Don Rubin (I think), pushing back against some of Adrian Raftery’s stuff on BMA (more succinctly, you can’t do BMA if you don’t know what you want to use it for).

    • I suspect this sort of wisdom tends to float around in part because of the popularity of time series methods in finance — where models are deployed in an fast-changing adversarial environment. A model which generalises well will be exploited, and the process of exploiting it will often stop it generalising well.

Leave a Reply

Your email address will not be published. Required fields are marked *