Dumb statistical models, always making people look bad

This is Jessica. In my Prediction for Decision-making class last quarter we talked about human predictions and decisions relative to algorithmic ones. This led to a discussion about why it’s often hard to demonstrate the value of human knowledge once you have a decent statistical model. 

Some readers may be aware of studies under the guise of “clinical versus statistical prediction,” most of which have found that statistical models generally outperform or match the accuracy of human predictors (see, e.g., Meehl’s 1954 book, where he summarizes the evidence at the time, and other meta-analyses. Then more recently (last 10 years or so) there has been an uptick in studies showing that when you give human decision-makers access to AI model predictions, they tend to do worse than the AI alone

There are a few ways to look at this from the standpoint of information that is available to the decision-maker. One is that human knowledge is valuable for guiding developing the model, but once you have a statistical model, it’s a better aggregator of the information. This is echoed by research on judgmental bootstrapping, where a statistical model trained on a human expert’s past judgments will tend to outperform that expert

This can be seen as a result of how we evaluate predictions. Regardless of how a human might arrive at some judgment on a specific case, we have to evaluate them actuarially (i.e., over a set of cases defined on some reference group, like other patients with similar characteristics). Minimizing loss over aggregates is what a statistical model is designed to do, so if you evaluate human judgment against statistical predictions in aggregate on data similar to what the model was trained on, then you should expect statistical prediction to win (unless your model really sucks) because it defines optimal use of the information for the task you’ve set up. From this perspective we shouldn’t be surprised. 

But the difficulty of showing the value of human expertise once you have a decent model can still make people uncomfortable, because it seems to contradict intuitions we have about what humans bring to decision scenarios. There are ways in which these intuitions can be shown to be misguided, but also some perspectives from which they seem to hold weight. 

Myths about human advantages

First let’s talk about what common beliefs about the superiority of clinical judgment tend to miss. For example, there’s an intuition that people can detect when a new instance is anomalous and step in to improve on the model’s prediction. I.e., there will always be longtail events that are underrepresented in training data, and people will know better what to do on these cases. A frequently cited example is what to expect from the person with the broken leg. In his book Meehl uses the example of estimating the probability that a professor, who is often at the movies on Tuesday night (with probability 0.9 according to the actuarial table) will go to the movies this Tuesday. Imagine that we know that he just broke his leg and got a hip cast (and let’s also assume this is pre-Americans with Disabilities Act). The clinician knows that this rules out the movies since he won’t fit in the seat with his bulky cast. 

But how often does this kind of scenario arise? Grove summarizes Meehl’s analysis of why this example does not vindicate clinical judgment:

First, the base rate of people breaking a leg is low, and so “broken leg cases” need not substantially decrease the accuracy of statistical predictions. Second, in the broken leg example, we have a highly reliable theory allowing us to predict clinically that Professor A has a very small probability of going to the movies; the theory rests on the physics of fitting a person with a hip cast into a 1954 movie seat. However, behavioral science theories are extremely seldom as well corroborated as those of physical mechanics. Third, the broken leg in Meehl’s example reduces the probability of movie attendance to zero, or some figure close to it. By contrast, when rare events occur in applied psychology, the event seldom guarantees that a person will, or will not, engage in the behavior of interest. Indeed, human behavior is so multidetermined that even unusual events typically change the probabilities modestly, or at most moderately. In sum, it is easy to see that “broken leg cases” exist and offer an opportunity for the clinician to do what the formula cannot. However, as Paul always maintained, it is very difficult to know whether a given case is a bona fide broken leg case, i.e., whether the clinician should overrule the actuarial prediction for a particular individual, or follow it.

A variation on the intuition that people can deal better with anomalies is that they can detect when a model is going outside its training distribution and adjust the prediction. But what do we have to assume for that to be a reliable benefit of human judgment? Larry Hedges, a fellow Northwestern faculty member, gave a talk awhile back where he walked through some examples from medicine and education. One was the example of an educator who has access to evidence from a randomized experiment (an estimate of the average treatment effect for some population) of some intervention that can improve student learning outcomes. The educator has more detailed information about the specific sample of students in their classroom relative to the general population from the controlled experiment. Should we trust them to be able to estimate the conditional average treatment effect if they apply the intervention in their classroom, and adjust how they apply it based on this? Hedges’ point was to show, given various ways of setting up this problem, that we repeatedly find that this would imply the educator had access to more data than is realistic, potentially many times as was used to estimate ATE from the randomized controlled trial.

This doesn’t necessarily mean that people can’t improve upon model predictions when a model is clearly out of domain (see, e.g., a recent paper on model-assisted decision making with Matt Hardy, Dan Goldstein, Jake Hofman, and Sam Zhang where we looked at how well people could predict weather in a new city given access to predictions from a model trained on a different city). We just can’t easily argue that this kind of tailoring of statistical evidence is due to people having more direct experience with the new domain. 

I remember Hedges concluding that if there’s anything that he expects to get from a human doctor over statistical predictions, it’s perhaps about knowing how to get him to adhere to treatment plans, e.g., reading his mood or what motivates him to pay attention. 

One of the other reasons that is sometimes cited for why humans should have some advantage is that they can act in the world to acquire more information as needed, while a predictive model, such as might be used to assist in diagnosing medical conditions or making treatment decisions, will not by default. But access to extra features alone can be hard to motivate as a general reason for expecting human superiority. If there are some important and easily elicitable pieces of information that aren’t available to the model (e.g., some info about the patient’s preferences, more recent measurements, or results from some test we order conditional on a risk prediction) probably we would be trying to either bring them into the prediction model or designing a separate statistical decision rule for incorporating them, and either of these would ultimately outperform the human judgment. So our reasons for having humans involved end up being because we haven’t yet done what we should be doing. 

Some better theories of human advantages

I think there is something to be said for human judgment along a few lines. 

One is about the ability to construct causal theories. This paper by Felin and Holweg on this perspective describes it as being not just a matter of differing information access between the human and model, it’s that the human can a) conceive of possibilities that may contradict the prior evidence, and b) construct tests to see if these “delusions” might hold water. Humans excel at forward-looking causal reasoning, which leads to opportunities for exploring novel ideas that the backward-looking imitative statistical learning paradigm can’t match. We can reason about counterfactuals and act when we think there’s something convincing.  They summarize this difference as humans being driven by data-belief asymmetries in ways that statistical models can’t be, i.e., people can hold beliefs that seem to contradict the prior evidence (e.g., the possibility of heavier-than-air human-powered flight immediately prior to the Wright brothers’ experiments) but which when explored through thoughtful experimentation hold weight.

I had a chat with Felin awhile back that got me thinking about the example of venture capitalists deciding what startups to back (which he mentions here). Apparently firms that use AI to help with such decisions are less likely to back the rare but truly innovative companies that go on to achieve major success compared to those that don’t. For this kind of decision under huge uncertainty and with built-in domain shift perhaps humans are not so bad.  

A related perspective to the humans-as-forward-looking-causal-reasoners hypothesis that I find compelling is that humans excel at identifying the right level of abstraction to reason about things in the world. There are plenty of demonstrations of statistical models relying on superficial clues (barns in the background of images of sheep) that are only loosely related to the target task (identify the animal) and sensitive to specific training conditions (conditions under which these particular photos were taken). Humans, at least superficially, seem more capable of developing abstractions that allow them to extract information that can be translated robustly from one situation to another; e.g., the relevant concept to reason about animals being near barns is farm  animals, not just sheep. Or imagine a self-driving car encountering an exploding fire hydrant. If such events are very sparse in the training data it will be very uncertain about how to proceed whereas a human will understand that it means water on the road and go from there.

P.S. It’s important to note that just because head-to-head comparisons with statistical models tend to make people look bad, this doesn’t mean that there isn’t sometimes value in human knowledge over statistical predictions. In AI-assisted decision-making research, there is work on learning to defer or learning with abstention, which is about optimizing performance by figuring out which decision tasks to give to the human versus the model based on estimated performance on different regions of the feature space. There have also been a few recent papers that attempt to quantify or test for unique complementary information on the part of the human decision-makers. We have a few papers about quantifying the value of complementary information in AI-assisted decision-making. Alur et al. have a related paper on tests for when human can discriminate between instances that are indistinguishable to a statistical model.  

These papers don’t get at exactly what it is that humans bring that is outside the known context of the problem, they simply show how to identify when humans have some complementary information to model predictions.

20 thoughts on “Dumb statistical models, always making people look bad

  1. The Meehl example is predicting an individual outcome. The papers look they like are comparing predictions of averaged outcomes. Then even averaging multiple averages (meta-analysis), and also the averaging of predictive skill of multiple humans rather than comparing the best/worst predictive skill.

    The vast majority of problems humans want solved regard individuals, not averages. Ie, is a prostate exam worthwhile for your specific case, not on average over the entire population.

    There is another aspect here where those averaged outcomes are also the output of statistical models. So it is kind of like asking whether statistical models are better at agreeing with other statistical models than humans.

    • >So it is kind of like asking whether statistical models are better at agreeing with other statistical models than humans.

      Yes, this is a good way to summarize the disparity between the original question of good individual decisions and what we end up evaluating.

  2. It is important to define what the relevant objective function for the judgment or decision is, too.

    If it is merely reducing MSE, this discussion is on-target. If, say, you want a judgment procedure that also produces a good explanation, a purely predictive neural net won’t do a good job (e.g., early XAI work). From a accuracy+explanation as objective fxn POV, human judgment offers much value — in part due to the abstraction and causal reasoning points you raise here. There’s lots of other kinds of values we get from human experts that we wouldn’t get from statistical models (e.g., social relations). So its worth explicitly scoping what we mean when we say “showing the value of human expertise once you have a decent model,” right?

  3. I think Kerem makes a good point regarding the relevant objective function for the judgment or decision. In insurance, adjusters are often asked to supply a case reserve, which is supposed to be an estimate of the ultimate settlement value of the claim. The adjuster typically has information that isn’t available in the data, but their estimate is influenced by a complicated incentive structure induced by the consumers of their estimates. For example, managers might perceive savings in claims that come in below their estimates, while actuaries are at least as concerned with the consistency of the estimates from the claims team as with their accuracy.

  4. Another issue often addressed with a second stage process is the cost of false positives relative to false negatives in a flagging pipeline. If the cost of false positives is substantially higher than false negatives, then attempts to identify some substantial number of those using human analysis can provide substantial utility.

    • Jessica –

      Humans, at least superficially, seem more capable of developing abstractions that allow them to extract information that can be translated robustly from one situation to another. .

      It strikes me that humans’ ability to develop abstractions, or maybe more accurately the type of abstractions they’re good at or prone to developing, is somewhat culturally influenced. I’m thinking of the well-known examples where East-Asians are more prone towards holistic, relational abstractions whereas Westerners tend towards discrete, analytic abstractions.

      What, if anything, might be lost by gathering all humans in the same basket here? (and for that matter, what might be lost by gathering all statistical models into the same basket)

      (Which recursively makes me wonder if there are any general differences in the different statistical models developed in different cultures🤔)

  5. “For this kind of decision under huge uncertainty and with built-in domain shift perhaps humans are not so bad.” That’s my expectation (or maybe hope?). In adversarial conditions, where the data are generated by someone who is trying to fool the model, it’s easy to find examples where an otherwise competent model falls flat on its face, even though a human would likely not be fooled. (see, e.g. the marines who tricked the person-detection algorithm by doing somersaults https://www.businessinsider.com/marines-fooled-darpa-robot-hiding-in-box-doing-somersaults-book-2023-1)

    On the other hand, adversarial data isn’t sufficient, since models can competently play some adversarial games.

  6. Part of why this “dumb models perform better than doctors” stuff feels unintuitive may be that we have a sense that the full scope of doctoring involves much more than the narrow predictive tasks where statistical models often excel. Just in the diagnosis space, we also need to assess the quality of the lab or imaging tests the patient has had, order better or different ones if needed and figure out how to reconcile discordant results, listen to the patient’s story and assess its accuracy, ask clarifying questions or call a family member or other doctor to get more info, etc.
    Yes, you can train models to do a lot of these things but there are so many different tasks, and there are other issues too: the available tests and their performance characteristics change based on medical advances, medical centers have different tests available, and some patients’ insurance won’t authorize certain tests or procedures. So we shouldn’t read too much into the models’ superior performance on a narrow diagnostic task.

    Great post by the way, I’m going to read all those references!

  7. One I have heard from farmers and fishers is that rural banks now make decisions about loans using models based on cities that can’t see the business logic of rural assets like specific pieces of land with specific soil and climate and management (they are great at loaning people with quota in the supply-management or fisheries system money to buy their neighbours’ quotas). Another is the examples discussed elsewhere on the blog where the model is pseudoscientific crap but you need people with both domain and stats knowledge to show that, and if you get rid of the experts that won’t happen. I think that Dan Kahneman insisted in his early work that the Israeli military accept some candidates who his models said should fail, so if they started to pass their courses that would be a signal that the model needed adjusting.

  8. I’m pretty ignorant of this whole field, but I’ve long been troubled by the emphasis on predictive accuracy as a metric for comparing machine and human decision-making. Of course, accuracy matters, but what I wonder is how we might compare them wrt perceiving the risk, as well as the cost, of error. Yes, humans are often victims of hubris, but a good human decision-maker has a sense of how much credence to give to a prediction, even beyond formal representation. Is this capacity exceeded by machine methods? And how would we know? Is there a literature on this topic?

  9. I think one very useful case of human intuition is context change, humans are quite good at detecting mode change based on exogenous factors, especially domain experts who constructed the model with some kind of insight into why the model should work to begin with. Statistical models are generally poor at understanding exogenous shifts and reacting to them, you can of course create a model to supervise the model to supervise the previous model and each will be more and more limited by the available data, humans on the other hand are very good at passive aggressive learning. If a model shows something it’s not very certain about it’s very unlikely it changes it’s mind 180 degrees while humans are inherently good at doing that under uncertainty.

    There are many domains usually in the context of imperfect information where a decent human supervisor with a decent model excel over amazing model as a result of the above is my speculation.

  10. I have a story that is only tangentially relevant, but I’ll tell it anyway because it’s related to several issues raised here, including the fact that maximizing “predictive accuracy” isn’t the right approach if mistakes in one direction are much worse than mistakes in another direction.

    A few years ago I was a physically fit 55-year-old who had never had high cholesterol, but my cholesterol numbers were gradually edging upwards year by year in spite of my continued healthy diet and exercise…and my dad had a quintuple bypass (the Full Letterman!) when he was around seventy due to clogged coronary arteries. Still, according to CDC recommendations I was not a candidate for taking a statin: statins were (and still are!) recommended for people that age only if they have LDL-C over 190 mg/dL, or have diabetes, hypertension, are smokers, or have dyslipidemia. So, no statin, no testing, no nuthin. But I asked for a referral to a cardiologist, and the doctor gave me one.

    The cardiologist said the same thing as my general practitioner: Nope, no need for a statin, no need for more testing. I pushed him on it a bit — hey doc, my dad had quintuple bypass surgery, my cholesterol has been increasing year by year, maybe I should do some additional testing or something? But he said “I follow the CDC recommendations and they don’t call for anything, so, no, I’m not prescribing anything.”

    Fortunately, my mom wanted more. She said a ‘cardiac calcium CT scan’ only costs about $600, and that I should get one. OK, sure mom, if it’ll make you feel better. So I got a cardiac calcium scan and it showed I have lots of calcification in my coronary arteries, including the Left Anterior Descending artery, a.k.a. The Widow-Maker. Of course I sent the results to my cardiologist, who immediately put me on a statin: the CDC recommendations did not recommend a statin OR a cardiac CT scan for someone with my numbers…until my numbers included the results of the CT scan. There’s plenty of data about the relationship between those calcium numbers and the risk of a heart attack, and the numbers are pretty scary! A lot less scary with a statin, fortunately, but still not great!

    My cardiologist wasn’t even a little bit sheepish about the whole thing: Hey, he’s following the CDC guidelines, that’s the ‘best available science’, of course it’s not perfect for every decision and some small fraction of people will wish they had done something else, but that’s Monday-morning quarterbacking, you have to look at the decision based on the information you knew at the time.

    I pointed out that I can look up the CDC recommendations for myself, so if the doctor is just going to follow the recommendations then what do I need him for, can’t I just go straight to the CDC and cut off the middle-man? He laughed, although I was not joking.

    I’m telling this story for a few reasons:
    1. If you are over about 40, at least consider getting a calcium CT scan. Do it if you can easily afford the $600, even if your doctor won’t prescribe it for you, and consider it even if it’s a stretch. If you’re over 55 or so, just do it.
    2. WTF is wrong with the CDC? There’s just no way that it makes sense to recommend no statin and no farther testing for someone with a family history of heart disease and a steadily increasing cholesterol number. You can do hundreds of calcium ct tests for the cost of treating one heart attack victim — assuming the victim survives.

    I’m not sure whether this is an example of a ‘simple model’ doing well or poorly, compared to human doctors. I don’t doubt that given the information available to the doctors, it was unlikely I would have an issue with my cardiac arteries: I’ve had regular physicals that never showed elevated cholesterol, I eat a healthy diet, I get a lot of exercise… based on the available numbers I was at low risk. The simple statistical model said I was at low risk, and the doctors’ gut feeling was also that I was at low risk.

    But what about the ‘value of information’ calculation? I think my mom had it right and the CDC has it wrong: even for someone with numbers like mine there was perhaps a 3% to 10% chance that the test would change the decision about whether to take a statin…which is what ended up happening.

    I’ll leave it to someone else to connect the dots to the original post.

Leave a Reply

Your email address will not be published. Required fields are marked *