Andreas Graefe sends along this paper (with Helmut Kuchenhoff, Veronika Stierle, and Bernhard Riedl) and writes:
We summarize prior evidence from the field of economic forecasting and find that the simple average was more accurate than Bayesian model averaging in three of four studies; on average, the error of BMA was 6% higher than the error of the simple average. We also add new evidence by reanalyzing the results from a study on election forecasting, which was published in Political Analysis and promoted the use of BMA for combining forecasts. However, the authors did not compare the results to the simple average as the widely established benchmark. We did: the error of BMA was 31% higher than the simple average.
I love this sort of empirical evaluation. Here are my quick thoughts:
One way this sort of “overfitting” finding can be interpreted is that you have prior information that is not being used in the model. If a simple average performs relatively well, another way of interpreting this is that you are fitting a model whose coefficients have very very strong prior information constraining them to be equal. My guess is that, in general, a slightly weaker prior would perform even better. But if the prior is too weak, you’re throwing away information and then you get the poor performance identified with BMA above. (BMA is only as good as the models it is averaging.)
And yet, out in the real world, all those million dollar data mining competitions are being won by ensemble models that average over models.
Weighted averages I should add
In Bayesian world those kind of ensemble models are more like Bayesian model compination rather than BMA (interesting article about BMC: https://axon.cs.byu.edu/papers/Kristine.ijcnn2011.pdf )
It seems that model combination implies a mixture model, where you could view each observation as being drawn from a different model in the mixture. Whereas model averaging implies the whole dataset is drawn from one of the models, just that we don’t know which. Would that be correct? If so perhaps it gives some intuition about the pros and cons.
EBMA (as used in this paper) sounds like one particular incarnation of all the possible ways the component models of an ensemble might be weighted. (I might be wrong)
Do the competition models use EBMA or something more sophisticated?
It’s usually linear regression or something similar. Ie. first use N models to make N predictions and then usese predictions as a input for linear regression (the sum of weights doesn’t have to be 1 in this case, but otherwise almost like weighted average). Something like that was used by Netflix winners (paper: https://www2.research.att.com/~volinsky/netflix/Bellkor2008.pdf ) and in many Kaggle competitions. In general this kind of approaches are called stacking in machine learning literature.
Idea behind BMC is pretty much the same (link above): first make predictions, then take some weighted combinations of those(=create new models) and finally apply BMA over those new models.
I’m trying to understand Andrew’s comment about the “very very strong prior” constraining coefficients to be equal and overfitting as a culprit for BMA’s poor performance.
If we have E(Y) = XB, with B in R^p, this prior is over B? So, say, B ~ Normal(b0, sd=.01)? I can see how this would reduce the variance of the coefficient estimates, but how does this relate to using a simple, unweighted average of models, which I interpret as implying a flat prior over the space of models?
I found this snippet from the paper highly non-intuitive:
If a model’s past performance is *unrelated* to future performance what’s the point behind any validation? Or even behind using any empirical data at all to train or tune a model?
Might as well build fully open loop models, guided only by fundamentals and theory?
What gives? What am I missing in my naive understanding?
I agree, I think the statement that past performance is unrelated to predictive performance, if interpreted as written, suggests that the whole enterprise (at least in terms of weighting models relative to each other) is a waste of time. I think this statement is probably a bit strong.
To me a big part of the issue comes down to the relationship between the quantities you are using to score models and the quantities you are trying to predict. If these quantities are basically the same (limiting case would be another sample from the exact same population as used to train the model(s), looking at the exact same outcome), then it might be reasonable to penalize quite harshly those models with poorer fit to the data (overfitting issues aside). On the other hand if the outcome we are interested in is very different to the outcome we are using to evaluate the competing models, then it is probably dangerous to penalize them too harshly.
This was my takeaway from arguments by Andrew and Don Rubin (I think), pushing back against some of Adrian Raftery’s stuff on BMA (more succinctly, you can’t do BMA if you don’t know what you want to use it for).
I suspect this sort of wisdom tends to float around in part because of the popularity of time series methods in finance — where models are deployed in an fast-changing adversarial environment. A model which generalises well will be exploited, and the process of exploiting it will often stop it generalising well.
The reason I was told that this guy https://en.wikipedia.org/wiki/Ferdinand_de_Saussure switched from economics to linguistics.
He thought that economics was the science of how people value things and if you communicated your studies on this to them, they would change how they value things.