This is Jessica. I was talking to a colleague yesterday about a paper I’m writing on prediction uncertainty quantification, and he commented that a statement I had proposed adding–that no model fit to finite data can achieve perfect calibration–was immoral. He suggested it would be misleading because (in his words), it is easy to achieve vanishing calibration error in theory and that is enough.
Putting aside for a moment his reasoning, I interpreted his reaction that it was immoral as saying that we don’t necessarily ever care about perfect calibration, we care about good enough calibration. That I agree with. But I was surprised that he labeled a statement of statistical fact “immoral.” The paper is largely about what to make of marginal coverage guarantees from uncertainty quantification in certain decision settings. In my view, pointing out that perfect calibration is not possible in practice is a useful starting point to this topic, because being explicit about what’s ideal versus possible helps make clear what we are really talking about when we say that the output of some model or method for quantifying uncertainty is “calibrated” or has some “valid coverage guarantee.” We’re usually actually implying something more complicated about some acceptable level of calibration. Similarly, it’s confusing when someone says that getting marginal coverage from an uncertainty quantification approach is not useful, because in reality we are always marginalizing over something. So it would seem like they actually mean something else but they are not communicating it well.
But I also appreciated his comment, because being too adamant about pointing out that we can’t obtain the ideal thing can be obfuscating, for example if it leads people to believe that the best existing approach we have for some goal is not worth adopting for reasons that are more about fundamental limitations on what we can learning from data than the specific method.
The fact that two people could agree on a practical point (that perfect calibration is not necessary) but feel differently about the appropriateness of pointing out a statistical fact got me thinking about the kind of dance we are often doing with language when we write about statistical methods, especially when we don’t expect the audience to know much about the associated theory. It’s important to be aware of what the theory does and does not establish when talking about a method, but it’s also complicated sometimes to bring it in, because the relevance of the facts to practice is not always obvious. The relationship between theory and practice itself can sometimes seem arbitrary. For instance, theory is usually the best source for understanding how we need to qualify statements we make about what we can achieve in practice from some method, whether we are talking about calibration, coverage or some other property where we can only hope to approximate what is optimal. But theory can also seem to completely sidestep other practical details that are critical in practice. In many applications, for example, no matter what we can say asymptotically about a method, we are stuck hoping our particular sample was not too extreme. Like my colleague’s comment about vanishing error, we end up talking about properties of repeated use rather than any particular use. It’s not that the theory tries to hide these realities; we should just expect the theory to focus on the properties of a statistical method that we can say something about with the tools at our disposal, even if these properties are not obviously solving the problems in practice.
We could say the difference between statistical factuality versus practicality is the difference between what is technically true and what is sufficient in practice. How we handle this when writing or talking about methods becomes a matter of art, because often we don’t study the difference directly. Typically we either care about what seems to work in practice (on the applied side) or what we can prove (on the theoretical side). And so individual authors have to decide how to move between these orientations. You could say that how one negotiates the gap between what is technically true and what is practically relevant is the “poetry” part of talking about statistics. It’s the process of figuring out how to bridge the gap between what is true or ideal about a method and what is important to our experience.
Asking whether it is morally acceptable or ethical to present a fact can be relevant because how we handle these kinds of dances is ultimately a function of our personal beliefs about what is valuable to the reader. What statistical facts you think are relevant and how you present them is a value judgment that depends on your personal frame of reference. For example, my colleague, who is a theorist, saying it’s easy to achieve vanishing calibration error might seem to him like he’s stating a simple fact, but from a practical standpoint one could label that statement as immoral or misleading because it fails to acknowledge that 1) he made it in the context of a discussion focused on what is true under finite data, not asymptotically, and 2) there’s a difference between the asymptotic result being true with respect to some simple (e.g., binary outcome) case and it being true about calibration in general.
Ultimately, I think this poetry part of writing about statistical methods is part of what makes the work seem important: there is some responsibility to the reader.
P.S. My use of the term poetry here is inspired somewhat by Kevin Munger’s recent blogging, though he uses it differently, distinguishing between the poetic truth of a statement like “social media is an echo chamber” versus its poetic validity, whether there is a valid mapping from the theory to the hypotheses to the observations from some study or set of studies.
Jessica:
I’m confused about something right at the beginning of your post, which is when you say you’re talking about finite data and then that there is “vanishing calibration error.” “Vanishing” is an asymptotic property, no? For finite data, I’d think that what’s relevant is the actual calibration error, not its behavior as N increases. But maybe I’m missing something?
As a separate matter, I can’t recall any colleague ever characterizing any of my published writing as “immoral,” so maybe the relevant variable here is not the content of your paper in question, but rather your colleague who expresses himself in such an unusual way!
Yes, his statement was not applicable to finite data. I interpreted it as his way of saying that my stressing the impossibility of perfect calibration was not helpful, but you’re right, it’s also confusing because he’s bringing in asymptotics in response to a statement about finite data.
On the morality question, apparently some theorists like to use term “morally correct” for things that are not correct, but close enough that it might as well be assumed to be correct. So I think that’s where his negative reaction to distinguishing perfect calibration from achievable calibration came from.
I interpreted “vanishing” as “vanishingly small.”
In my experience with engineering modeling, calibration error, and measurement error more generally, always played a smaller role in what the modelers were doing than outsiders expected. If you asked someone on the periphery of the project what they thought was the most likely scenario for the model going off the rails, they would say something like “well the flowmeter is not very accurate” or something like that. But all that stuff had been worked out. When the model REALLY went off the rails it was from one of two root causes, either we bumped into a heretofore unknown edge of the operating envelope, or the modelers simply failed to anticipate an inevitable change (I call this latter outcome “poverty of the imagination).
So maybe this:
From the perspective of the modeling goals, calibration error generally plays a vanishingly small role, so it is “immoral” to get people to focus on it any more than they already do.
When you say “calibration error generally plays a vanishingly small role,” I think you are saying that the difference between the model’s calibration on finite data and its calibration on infinite data from the same source is often not important. But when there is some bigger change in the conditions when we go to apply the model (which is one way of interpreting what you call bumping into a heretofore unknown edge of the operating envelope, which is also a kind of calibration problem with the original model) then it can matter.
I guess my reasoning that its good to point out that calibration is never perfect is partly because we’re often trading off one type of error for another. For example, one approach to deal with potentially biased training data in ML is to get some new data and post-hoc calibrate the model to that. But in that scenario you may be using a lot less data than you trained the model on, so your attempt to overcome bad calibration from biases in training data puts you in a situation where you may actually have to worry about the first kind of calibration error.
But all this also depends on how it gets used, so it comes down to what sort of assumptions you are implicitly making about what matters in the practical application.
Your point, “Asking whether it is morally acceptable or ethical to present a fact can be relevant because how we handle these kinds of dances is ultimately a function of our personal beliefs about what is valuable to the reader. What statistical facts you think are relevant and how you present them is a value judgment that depends on your personal frame of reference.” is extremely important.
My colleagues and I have created an ethics review process for AI models at our company. Our process requires teams so submit a substantial amount of testing data, sliced many different ways. We get push back from teams that would rather submit a fraction of the data they deem relevant.
Your statement will now be my official reply. Thanks!
Chris:
This reminds me of my article, Ethics and the statistical use of prior information.
The bigger problem is that the world is not an RNG and the goal of Bayesian models is not to be “calibrated” but rather to provide a good measure of how uncertain outcomes may be.
If I tell you there is a 80% chance of rain for 100 days in a row and it turns out only 70% of those days have rain where you are is this “miscalibration”? Is the goal of your Bayesian weather model to provide a frequency estimate? Or is it to provide a method of making decisions about crop planting, construction projects, reservoir safety and power generation, and sporting events and such like that?
Most times it’s the latter, and most times having frequency properties is not even well defined. There is no long run correct frequency of rain. (Yes/no) There isn’t even a long run correct frequency distribution of rain depth (inches or cm). Perhaps you’ve heard of climate change, ElNino, desertification, deforestation etc etc.
Frequency Calibration is the wrong way to think. In this sense it’s “immoral” because it’s misleading in the way that misinformation about climate change is misleading. Moving the conversation away from reality and towards a broken view of science.
Note, there’s a major difference between situations in which there are a preexisting set of measurable things at a point in time, not all of which are measured, and “all things that might happen in the future”
Consider the difference between the distribution of breaking strength of suchandsuch brand bolts being held at some particular Home Depot on some particular day… A well defined finite population of finite values. And “the distribution of future bank account balances at Wells Fargo bank of people living in CA”
There can be no such distribution of the bank balances because inflation ensures that more and more of the balances through time will be bigger and bigger values, furthermore it’s entirely plausible that Wells Fargo merged with Chase or some such thing to become a different bank entirely, or that nuclear Armageddon causes WF and dollars both to cease to exist…
Lots of things have no well defined frequency distribution. It’s better to model the future as a dynamic process than as repeated samples.
Isn’t there a middle ground? There is such a thing as the distribution of WF bank balances if the future unspecified processes continue as in the past. This distribution would exhibit the variability based on past patterns, possibly based on a number of measurable factors that have influenced these balances in the past. The future “dynamic process” might contain any number of new events, some of which may essentially be “black swans.” While the frequency distribution based on past balances will not reflect these new factors, they can provide a benchmark for a more exploratory analysis (perhaps simulations) of how future balances may differ from the past. I don’t see why modeling the future as a dynamic process and frequency analysis based on repeated samples are mutually exclusive.
Dale, definitely frequency can be a thing you model with Bayesian probability. But you need to keep straight in your mind what is what.
For example, suppose I break the future into years. I could say that within a year the frequency distribution of balances is reasonably well defined… What I mean perhaps is if I pick a random person and a random day I have an end of day balance, and there are a finite number of such things and so a frequency distribution is defined automatically.
Now, how much do I know about it’s shape? I need to define Bayesian probabilities over parameters that define the shape. And supposing I have ending balances for 2500 people on various days in 2023 I don’t know a whole lot about the shape… But I know something. So perhaps I can define a gaussian mixture model for the frequency, and I define a Bayesian probability over the location and scale and mixture fraction for each mixture component.
Now I have a whole family of PDF curves I could draw representing my uncertainty about the frequency.
Totally 100% legit modeling exercise. But not at all similar to what frequentist do in practice as far as I know.
>Frequency Calibration is the wrong way to think. In this sense it’s “immoral” because it’s misleading in the way that misinformation about climate change is misleading. Moving the conversation away from reality and towards a broken view of science.
I am more inclined to agree with this view on immorality than my colleagues; thanks for pointing out the implicit frequentism. But can you say more about your definition of frequency calibration? Would you say the concept of calibration itself is problematic, because frequency is always implied in some sense? Or are you arguing that calibration only makes sense in a sequential framework? (relatedly, I’m curious what you think of Dawid’s view of the well-calibrated Bayesian)
Oh, I just saw your above response to Dale which helps answer some of these questions.
Great. Let me know if there’s something specific still confusing you’d like me to elaborate.