Jay Kadane sends along this this new paper on subjectivity and objectivity in Bayesian statistics, which of course reminded me of my 2017 paper with Christian Hennig on the same topic.
In response to my article, Jay wrote that he agreed with much of what we wrote, but:
My [Jay’s] qualm is the emphasis on consensus, which in my view, is not necessarily a virtue at all. Let me give an example from my own work. In Grossman et. al. (2011) [Bayesian Analysis 6 (#5), 547-572, doi 10.1214/11-BA62], we address the question of whether the number of hurricanes in the North Atlantic had increased in the last 150 years (because, perhaps, of the industrial revolution and the resultant increase of carbon in the atmosphere.) A reasonable question, I think.
It turned out the there’s a standard database of Hurricanes (HURDAT) run by physicists, who require at least 2 independent measurements of wind speeds greater than some threshold I forget, for a storm to be listed The issue we found is that our ability to detect hurricanes has increased enormously in that period. In the beginning, there were landfall and ship observations only. At the end of the period there were satellites that see everything. So we had to model the probability of observation as a function of time. When we did that, we found that by changing the prior on one parameter in seemingly innocuous ways, we could get increasing, constant or decreasing frequency of hurricanes. So our bottom line is that we don’t know, and we expect nobody’s going to know, the answer to the question we posed. Pressure to achieve consensus here is counterproductive. Sometimes we just have to live with not having a consensus, not knowing what we wish we did. Sometimes lack of consensus is the story, and it would be wrong to make this non-virtuous.
Fair enough, and indeed Jay’s example demonstrates the virtue of “investigation of stability,” listed in table 1 of my article with Hennig:
Regarding consensus . . . I think that in Jay’s example, the appropriate consensus is that we just don’t know! There are some settings where I feel we just don’t know, but there is no consensus on that (see for example here), but for this hurricanes thing, perhaps a good statistical analysis accounting for uncertainty can help to create such a consensus agreement of ignorance. Scientific consensus can include uncertainty.
I agree, though, that a lot of people–even in fields such as computer science or finance where uncertainty is prevalent–often seem to want a consensus of certainty, even when it is not possible. This came up last year in the context of election forecasts. Back in 2024, for our Economist magazine election model, we put together economic and political fundamentals and state and national polls and came to the conclusion that the presidential election was likely to be very close, and the two candidates each had about a 50% chance of winning. Lots of people understand this idea, that a careful forecast can end up giving probabilities close to 50%, but some people were confused on the point. One person faulted our forecast for not predicting a “landslide,” which was kind of funny given that the election was not a landslide at all; it was very close. Another person characterized the 50-50 forecast as “giving up,” which was completely wrong because to say that you think the election will be close is itself informative. It’s similar to a football game where the point spread is zero: this represents information on the relative strengths of the teams.
Anyway, my point here is that in 2024 there was a consensus among reasonable forecasters that the election was highly uncertain–some forecasters gave Trump a 2/3 chance, which was also reasonable, but that would be considered close too–, but there were some people who couldn’t handle this uncertainty and didn’t recognize it as a consensus. So I see Jay’s point.

This got me thinking whether you could argue for a negative value of information in cases where a complete lack of certainty leads to better decision-making. E.g., the small disutility of carrying an umbrella if it didn’t rain would still be vastly preferable to the disutility of wearing your favourite white linen shirt and then having to walk to work in a downpour.
Normally, VoI is positive, but (incorrect) consensus would be worth substantially less than 0 if it leads to worse decisions because of misallocated resources. You could just flip the sign, but bad consensus < 0.
Robin:
Related, perhaps, is that if you do inference with a normal/normal or beta/binomial model, your posterior variance is always lower than your prior variance (this is a homework problem in BDA!), i.e., getting more information always causes the variance to go down.
More generally, the posterior variance can sometimes be higher than the prior variance and sometimes lower, but if the model is correct (that is, averaging over the prior), the posterior variance will be lower than the prior variance in expectation. I think that’s a homework problem in BDA too.
If the model is not normal/normal or beta/binomial or some other simple cases, and the model is wrong, then all bets are off.
Consensus around the wrong model is a nice explanation. I really should look through BDA once I finish with Statistical Rethinking.
Typo:
“Lots of people go this, but there was some confusion.”
Fixed; thanks.
A bit of whiplash for me.
A few columns back, someone posted a quote from a guy who said he can just change his prior and get the answer he wants from the data. Andrew responded that he “hated” this framing because the data model can be tweaked just as easily. In the scenario that Andrew described, the priors were guided by empirical evidence while the data model was the creation of the scientist doing the research. So Andrew’s comment made sense.
But I couldn’t help noticing that the scientific scenario being discussed was different from the way Andrew framed it. In climate science, the data model consists of a set of physical relationships taken from atmospheric physics, with very little room for researcher degrees of freedom. It is plug-and-chug. Meanwhile, with some of the basic mechanisms of climate such as the water vapor cycle still being poorly understood, there is very little empirical evidence that can be used to guide a researcher to an even partially objective prior to plug in to the existing data model. The result is that in fields like extreme climate event attribution, the researchers just tweak the priors until they get a sort of Goldilocks posterior, big enough to look scary but not so big as to look absurd. They cannot tweak the data model, that would be way above their pay grade.
As far as I can recall, this is the first post on this blog that disavows an ability to generate a meaningful result using a Bayesian method under deep uncertainty, with the idea having been that a Bayesian can always do SOMETHING that is better than nothing. Certainly, Anoneuoid has embraced this view, going so far as to claim that the worst Bayesian analysis still has to be better than the best frequentist analysis.
Matt:
You write, “In climate science, the data model consists of a set of physical relationships taken from atmospheric physics, with very little room for researcher degrees of freedom.” Unfortunately, no, that’s not the case! I’ve done some climate science–tree-ring reconstruction; see here–and the data model, which links the data to the underlying parameters, had to be constructed, and there’s lots of uncertainty about it.
Regarding your last paragraph: there have been other times I’ve said that I don’t see exactly how Bayesian inference can help in a problem. See here for an example. I still like the idea of Bayesian inference for that particular problem in principle, but the details leave me stymied.
Hi Andy,
What an interesting set of comments! But I have qualms about your solution. When two (or several) analyses disagree, your solution is to conclude that “we” don’t know (and therefore agree that we don’t know). But this can hide real and serious differences of opinion.
I am for truthful statements of (posterior) opinion. If they coincide, well and good . And “virtuous” according to you. But if they do not coincide, there is virtue in having the opinions clearly stated, along with the rationale for each opinion.
Sometimes we have to tolerate that we disagree, and we can do that without being disagreeable.
All the best,
Jay
Jay:
Sure, what you’re saying corresponds to V5 and V7a in my list above. I’m not saying that consensus is the only virtue.
“There are some settings where I feel we just don’t know, but there is no consensus on that”.
In a finance context…
The current price of the SP500 market index is a consensus measure of the SP500 in a month.
Uncertainty about the consensus, about where the market will be in a month, is measured by the VIX. The VIX is, very roughly, the market’s estimate of the standard deviation of the SP500 in a month. It is based on SP500 option prices. Low VIX says the future price will be similar to the present. High VIX says a big range for the future price.
What will the VIX be in a month? Uncertainty of uncertainty. It is measured by the VVIX, which is based on VIX options.
Says the Market:
High VIX, High VVIX: Future prices will be in a wide range, but there is no consensus on that.
High VIX, Low VVIX. Everyone knows there is a wide range about future prices.
Low VIX, High VVIX: The future will be like the present, maybe.
Low VIX, Low VVIX: The future will be like the present, likely.
As co-author of the cited 2017 paper, the consensus “virtue” was and is very important to me. I believe that striving for consensus, or, in other words, for generating a world view on which as many people as possible can agree *in free exchange*, is a major aim of science. This has to do with my constructivist philosophical leanings. One could hold that science is interested in finding out the truth, and as long as it does so, consensus is not really relevant. But as human beings we don’t have unmediated access to the truth, and an important way our ideas about reality can be corroborated is the agreement of others (the relevance of agreement obviously correlated with the amount of relevant knowledge and understanding that the person has whose agreement is in question). I hope most agree that the scientist’s job is not only to discover, but also to convince, and furthermore to adapt when confronted with justified criticism. The “Consensus” item in our list of virtues is about striving to enable consensus regarding our results and discoveries, and to integrate them into the accepted scientific body of knowledge. What we do there is that we acknowledge science as an essentially social endeavour.
Unfortunately this point is often misunderstood as embracing “pressure” to achieve consensus, as mentioned by Jay Kadane. My view is that it’s really rather the opposite. Scientific consensus should be achieved (as far as it can be achieved) based on free exchange, and the best scientific consensus is the one that remains stable in the face of criticism and opposition, taking all the points raised into account. So indeed striving for consensus would encourage criticism and opposition in order to achieve a more general and stable consensus. This also implies that science needs to engage with uncertainty and with different and even opposing points of view. The formula “let’s agree to disagree” makes some sense as agreement on uncertainty and the existence of divergent legitimate points of view is certainly more constructive than if one side tries to impose their views on the other, or some pressure is exercised to reach whatever consensus. But the key issue is that the importance of consensus in my view is about communicating science in a way that facilitates reaching consensus (see the subitems for what this can involve), and about acknowledging the key role of exchange and communication in science. It is not about saying that reaching consensus in itself is desirable if getting there involves pressure and ignorance or suppression of opposing views, i.e., anything that sabotages science in other respects.
If I understand correctly, we can (almost) never be fully confident that we have discovered an immutable truth, and if others arrive at the same conclusion independently, or dependently but from a contrary starting position, then we can be more confident taking the findings as true enough for decision-making.
Perhaps the tension here with your (quite reasonable) point underpinning Jay’s critique is that scientists stand to gain personally from establishing a consensus position regardless of accuracy? I’m reminded of Gigerenzer’s recollection of Kahneman’s belief that being perceived to be a leader in the field was more important than being correct.
Typo:
The citation for the Grossman et al. paper is quoted as [Bayesian Analysis 6 (#5), 547-572, doi 10.1214/11-BA62].
It should be [Bayesian Analysis 6 (#4), 547-572, doi 10.1214/11-BA621]. (If you Google 10.1214/11-BA62 it just links back to this page)
Jim Joyce’s “How Probabilities Reflect Evidence” provides a relevant viewpoint for this discussion. He distinguishes three ways in which probabilities reflect evidence:
– The overall balance of the evidence is a matter of how decisively the data tells in favor of the proposition. This is what individual probability values reflect.
– The weight of the evidence is a matter of the gross amount of relevant data available. It is reflected in the concentration and stability of probabilities in the face of changing information.
– The specificity of the evidence is a matter of the degree to which the data discriminates the truth of the proposition from that of alternatives. It is reflected in the spread of probability values across a credal state.
A quick example to illustrate: You have a misshapen coin and are interested in its bias. You judge the range of admissible priors that fit your evidence to be beta(0.5,0.5) through beta(10,10). For beta(0.5,0.5), you calculate P(0.4<p<0.6)=0.128. For beta(10,10), you calculate P(0.4<p<0.6)=0.651. The interval [0.128,0.651] represents your credal state concerning 0.4<p<0.6. You flip the coin 30 times and obtain 17 heads & 13 tails. You update your priors. For beta(0.5+17,0.5+13), P(0.4<p<0.6)=0.617. For beta(11+17,11+13), P(0.4<p<0.6)=0.791. The interval [0.617,0.791] represents you new credal state concerning 0.4<p<0.6.
Before the data, your credal state interval is wide, indicating a lack of specificity of the evidence. It is difficult to judge if the evidence supports the idea that 0.4<p<0.6, and the weight of evidence is low. After the data, the credal state interval is much more narrow, indicating increased specificity of evidence. The balance is now slightly in favor of 0.4<p<0.6 over its negation, and the weight of evidence has increased (another single coin flip won't influence the results much now).
My view is that epistemic probabilities arise entirely from judgments about what the evidence means, and they should faithfully reflect that evidence. This way of thinking ties everything together and lets us talk about probabilities purely through the lens of evidential judgment. Epistemic probabilities are subjective in the sense that they involve judgment, yet objective in the sense that these judgments are grounded in evidence. When the evidence is highly specific and its weight is large (for example, when a misshapen coin has been flipped many times), we should expect broad consensus about the balance of evidence.