Defining decisions in studies of visualization and human centered AI

This is Jessica. A few years ago, Dimara and Stasko published a paper pointing to the lack of decision tasks in evaluation of visualization research, where it’s common to talk about decision-making, but then to ask simpler perceptual style questions in the study you run. A few years earlier I had pointed to the same irony when taking stock of empirical research on visualizing uncertainty, where despite frequent mention of “better decision-making” as the objective for visualizing uncertainty, few well-defined decision tasks are studied, and instead most studies evaluate how well people can read data from a chart and how confident they report feeling about their answers to the task. Colloquially, I’ve heard a decision task described as a choice between alternatives, or a choice where the stakes are high, but neither of these isolates a clear set of assumptions. 

Then there is all the research being produced on “AI-advised decisions” – how people make decisions with the aid of AI and ML models–which has become much more popular in the last five years. Some of this research is invested in isolating decision tasks to build general understanding about how people use model predictions. E.g., according to one recent survey of empirical human-subjects studies on using AI to augment human decisions, these studies are “necessary to evaluate the effectiveness of AI technologies in assisting decision making, but also to form a foundational understanding of how people interact with AI to make decisions.” The body of empirical work on AI-advised decisions as a whole is thought to allow us to “develop a rigorous science of humanAI decision-making”; or “to assess that trust exists between human-users and AI-embedded systems in decision making”; etc. Reading these kinds of statements, it would seem there must be some consensus on what a decision is and how decisions are a distinct form of human behavior compared to other tasks that are not considered decisions. 

But looking at the range of tasks that get filed under “studying decision making” in both of these areas, it’s not very clear what the common task structure is that makes these studies about decision-making. For some, the point seems to be to compare human decisions to a definition of rational behavior, but sometimes this is defined by the researchers but other times it’s not. Sometimes the point is to study “subjective decisions,” like helping a friend decide if the list price of a house matches its valuation where the participants are intended to use their own judgment in thinking about prioritizing them.  

If we are going to isolate decision-making as an important class of behavior when we study interfaces, I think we should be able to give a definition of what that means. I get that human decision-making might appear to require hard to formalize sometimes (because if it wasn’t, why haven’t the humans figured out how to automate it?) And decision theory isn’t necessarily familiar if you’re coming from a computer science background. But it seems hard to make progress or learn from a body of empirical work on decisions if we can’t say exactly what classifies a task as decision  The survey papers I’ve seen on decision making in visualization or in human-centered AI conclude that we need a more coherent definition of decision, but no one appears to be suggesting anything concrete. 

So here’s a proposal for one way to understand what constitutes a decision problem: a decision task involves the person choosing a response from some set of possible responses, where we know that the quality of the response depends on the realization of a state of the world which is uncertain at the time of the decision. Additionally, for the fields I’m talking about above, we generally want to assume that there is some information (or signal) available to the decision maker when they make their decision which is correlated with the uncertain state. 

We can summarize this by saying when we talk about people making a decision we should be able to point to: 

  1. An uncertain state of the world. E.g., if we are taking recidivism prediction as the decision task, the uncertain state is whether the person will commit another crime after being released, which can take the value 0 (no) or 1 (yes).  
  2. An action space from which the decision maker chooses a response. Above I called the action space a choice of response, because I’ve seen a colloquial understanding of decision that assumes that a task has to involve choosing between some actions that correspond to something that feels like a real-world choice between a small number of options. It doesn’t. The response to the decision problem it could be a probability forecast or some other numeric response. What matters is that we have a way to evaluate the quality of that response that accounts for the realization of the state.
  3. A (scoring) rule that assigns some quality score to each chosen action so we can evaluate the decision. I think sometimes people hear ‘scoring rule’ and they think it has to be a proper scoring rule, or at least a continuous function, or something like that. It can refer to any way in which we assign value to different responses to the task. However, we should acknowledge that we can use scoring rules for different purposes even in a single study, and think about why we do this. E.g., sometimes we might have one scoring rule that we use to incentivize participants in our study (which may just be a flat reward scheme regardless of the quality of your responses, which is used a lot at least in visualization research). Then we have some different scoring rule that we use to evaluate their responses, like evaluating the accuracy of the responses. If we want to conclude that we have learned about the quality of human decisions, it’s the latter we care about more, but in many situations we should also be thinking carefully about the former as well, so that participants understand the decision problem the same way that we do when we analyze it . So we should be able to identify both when running a decision study. 

 And, optionally for studies comparing different interfaces to aid a decision maker:

  1. Some signal (or set of signals) that inform about the uncertain state that the decision-maker has access to in making their decision. This could be a visualization of some relevant data, or the prediction made by a model on some instance.   

This may seem obvious to some, as I am essentially just describing components of statistical decision theory. But I think making these aspects of a decision problem explicit when talking about decisions in visualization and HCAI research would already be a step forward. It would at least give us a list of things to make sure we can identify if we are trying to understand a decision task. And it could help us realize the limitations on how much we can really say about decision-making when we can’t specify all these components. For example, if there’s no uncertain state on which the quality of someone’s response to the task depends, what’s the point of trying to evaluate different decision strategies? Or if we can’t describe how the signals our interface is providing compare in terms of conveying information about the uncertain state, then how can we evaluate different approaches to presenting them?

Something else I’ve noticed is papers that refer to a decision task but then bury the information about the nature of the scoring rule that is used, especially that used to incentivize the participants, as if it doesn’t matter for anything. But if we give no thought to what to tell study participants about how to make a decision well, then we should be careful about using the results of the study to talk about making better or worse decisions – they might have been trying different things, doing whatever seemed quickest, etc. 

Also, some studies resist defining decision quality even in assessing responses, as if this would take away from the realism of the task. I think there’s a temptation in studying these topics to assume there’s some inherent blackbox nature to how people make decisions or use visualizations that absolves us as researchers of having to try to formalize anything. It isn’t necessarily wrong to study tasks where we can’t say what exactly would constitute a better decision in some real world decision pipeline, but if we want to work toward a general understanding of human decision making with AIs or with visualizations through controlled empirical experiments, we should study scenarios that we can fully understand. Otherwise we can make observations perhaps, but not value judgments.  

Related to this, I think adhering to this definition of decision would make it easier for researchers to tell the difference between normative decision studies and descriptive or exploratory ones. If the goal is simply to understand how people approach some kind of decision task, or what they need, like this example of how child welfare workers screen cases, then it doesn’t necessarily matter if we can’t say what the scoring rule/rules is – maybe that’s part of what we’re trying to learn. But I would argue that whenever we want to conclude something about decision quality, we should be able to describe and motivate the scoring rule(s) we’re using. I’ve come across papers in both of the areas I mentioned that seem to confuse the two. I also have mixed feelings about labeling some decisions as “subjective,” though I understand the motivation behind trying to distinguish the more formally defined tasks from those that seem underspecified. There’s a risk of “subjective” making it sound like it’s possible to have a decision for which there really is no scoring rule, implicit or not, but I don’t think that makes sense.

Of course there is lots more that could be said about all this. For example, if you’re going to study some decision task with the goal of producing new knowledge about human decision making in general, I think you should be able to go further than just specifying your decision problem: you should also understand its properties and motivate why they are important. I find that often there is an iterative process of specifying the decision task for an experiment – you take a stab at specifying something you think might work, then you attempt to understand how well it “works” for the purposes of your experiment. This process can be very opaque if you don’t have a good sense of what properties matter for the kind of claim you hope to make from your results. I have some recent work where we lay out an approach to evaluating the decision experiments themselves, but will leave that for a follow-up post.

P.S. These thoughts are preliminary, and I welcome feedback from those studying decisions in the areas I’ve mentioned, or other domains.

18 thoughts on “Defining decisions in studies of visualization and human centered AI

  1. Jessica:

    My quick suggestion is to take a look at the judgment and decision making literature starting from the 1960s. (The stuff in the 1940s and 1950s is interesting but will seem too obvious now to be useful.) For example, a standard way to handle decision making under uncertainty is to first consider decision making under certainty and then add the uncertain element to it. So I disagree with your item #1 that there needs to be uncertainty. Even for certain outcomes there can still be decisions needed. These can be thought of directly as decisions under certainty or else as conditional decisions in an uncertain setting.

    • Thanks. I guess it’s hard for me to imagine studying the effects of different interfaces (model explanations, visualizations, whatever) if the state is deterministic – what information does the interface then convey if not the state? Relatedly, I am reminded of this definition of the idea of “trust” in a predictive model as requiring some uncertainty hence vulnerability on the part of the human decision maker: https://dl.acm.org/doi/pdf/10.1145/3442188.3445923

      But what you’re saying makes sense, that you can still have a decision without uncertainty. So its something I need to think about more if I ever decide to write something more formal up about how to define decision wrt to vis/human-centered AI.

      • There’s at least one book on the topic: Decision Making Under Certainty, by Arthur Schleifer and David Bell. It’s part of a 4-book series of small textbooks. One of the others is called Decision Making Under Uncertainty.

        • It could be cast as a decision, where the uncertain state is true sentiment, action space is binary, signal is high dimensional (text).

          One way in which the view I’m taking here may differ from what Andrew is saying is that for the goals I have in mind, whenever you run an experiment that involves some variation in the state you can treat it as uncertain in any given trial a priori, even if it is fully determined by the signal. The reason we might want to do this is to understand, for example, the value of the information we are giving people in each trial for making the decision, relative to expected performance if they knew only the prior over states (e.g., the prior probability that the true sentiment will be positive over the trials in the experiment). If the signals have a low value of information for the decision task, that’s a bad experiment design. I have a recent paper on this kind of analysis for visualization experiments which I will post about soon. My coauthors (Ziyang Guo, Yifan Wu, Jason Hartline) and I are currently talking about how this can be applied to experiments on tasks like the one you mention in explainable AI research, which might be of interest to you.

      • Jessica:

        This discussion points to some issues that I remember thinking about many years ago when I used to teach a class in decision analysis. The recommended way we set up a decision problem is to start with goals, then move on to plans, then set up specific decision options and draw the tree of possible outcomes (which could be complicated if there are intermediate decisions along the way), then assign utilities to all the possible endpoints (the leaves of the tree), then make a decision plan starting from the leaves and work your way back to the base of the tree. The tree itself has leaves, uncertainty nodes, and decision nodes, and the rule is to average over the uncertainty nodes and optimize over the decision nodes.

        From that perspective, decision making under certainty might seem trivial: there are no uncertainty nodes so you just optimize over the decision nodes, which reduces to choosing the path of decisions that leads to the leaf with highest utility. One major challenge comes before the tree gets drawn, in deciding on goals, plans, and options at each node. The other major challenge comes in assigning the utility or value to the leaves. Even if there are no uncertainty nodes, it’s still a big effort to map complicated outcomes into scalar utilities, and I think that is most of what the “decision making under uncertainty” literature is all about. I think this connects to what you’re saying in that, even if the decision tree has no uncertainty nodes, you still have uncertainty as to how to value the leaves.

        The other thing I’ve thought about for awhile is that there could be a logic to adding uncertainty to the decision nodes. The idea is that even if it’s your own future decision, realistically you won’t know ahead of time exactly how you will act conditional on new information that comes in, so you can make a decision in part by averaging (rather than optimizing) over your own path through the tree. An appealing aspect of this approach is that it would simplify the math, because if all the nodes are uncertainty nodes, then evaluation of the tree is just a set of nested averaging, which is relatively easy to do (a big integral or a big simulation), as compared with the great difficulty of alternately averaging and optimizing.

  2. In a different context (economics) I’ve also struggled to find a useful formal definition of ‘decision process’. One resource I’ve found useful in thinking about this, though it is far from my field, is the work of the management theorist Henry Mintzberg. He provided some formal definitions of a ‘decision process’ – in a 1976 paper, ‘The Structure of “Unstructured” Decision Processes’, and his book ‘The Structure of Organizations’. As that title suggests, his focus is on decision-making by organizations, rather than individuals.

    Mintzberg distinguishes between three ‘phases’ of decision making – drawing on Simon’s Intelligence/Design/Choice framework, described by Dimara and Stasko – and seven ‘routines’, or fundamentally different types of activity that take place in a decision process. The (1) recognition and (2) diagnosis routines take place in the ‘identification’ phase; (3) search (to find ready-made solutions) and (4) design (to develop custom solutions) take place in the ‘development of solutions’ phase’; and the (5) screening of ready-made solutions, (6) evaluation-choice of custom solutions, and (7) ‘authorization’ routines take place in the ‘selection’ phase. As I understand it, the framework was developed based on empirical studies of ‘strategic’ decision making, but it is applicable to more routine operating decisions as well (these would involve less diagnosis, less design of custom solutions, and more search for ready-made solutions).

    This might be quite far from the problem you’re interested in, but if nothing else maybe the 1976 paper has some useful references.

  3. Do practitioners, especially the sort making critical descisions, actually use this theoretical framework of descision making much? Lot of it has been around since the 1960s but doesn’t seem to have made it to actual practice.

    • Are you asking if statistical decision theory gets used in practice, or something more specific about human decision making? It seems obvious that this style of modeling a decision problem gets used often in practice e.g., to optimize sequential decision processes in industrial and medical applications. If you mean do organizations go through the trouble of modeling and evaluating decisions that are ultimately made by humans working with some kind of data-driven displays, I don’t know of many examples. But my point in this post is not that all decision makers should try to understand their decisions this way, it is that if we as researchers are going to try to systematically study these scenarios, we need to pinpoint what problem we are studying, and without a more formal treatment of a decision problem it seems very hard to accumulate knowledge.

  4. Jessica, at your uni, James Glynn (I think), may appreciate a call from you… “when taking stock of empirical research on visualizing uncertainty”.
    *

    “Computer simulation provides 4,000 scenarios for a climate turnaround”

    “Tackling climate change requires numerous political, economic and social decisions to be made. However, these decisions are fraught with uncertainty…

    “This calls for sophisticated data analysis and visualization techniques,” adds co-author James Glynn, head of the Energy Systems Modeling Program at Columbia University in the US. The final file contains 700 gigabytes of data. The paper on this research has now been published in the journal Energy Policy.

    https://phys.org/news/2023-06-simulation-scenarios-climate-turnaround.html

  5. In their book Decision Analysis for Healthcare Management, Alemi and Gustafson attempt to define a decision. This is what they say in Chapter 1:

    ” A decision is made when a course of action is selected among alternatives. A decision has the following five components:
    1. Multiple alternatives or options are available.
    2. Each alternative leads to a series of consequences.
    3. The decision maker is uncertain about what might happen.
    4. The decision maker has different preferences about outcomes associated with various consequences.
    5. A decision involves choosing among uncertain outcomes with different values”

    I found the book referenced in Decision Synthesis by Watson and Buede. There are two interesting things that come to mind as I read your definition and theirs:
    (1) Do you evaluate a response based on information known at the time of choosing the response or after? And which of the two does the scoring rule capture? Does it measure the quality of the decision based on information known at the time of making the decision? In this case, the information available could be the decision-maker’s value function (including their own and other constituencies’ values), information that might correlate with future outcomes (for example online reviews if the decision is a purchase decision), and the measure of quality could be consistency of chosen action with information and the value function. On the other hand, if the decision is to be evaluated based on the outcomes then there is a difference in what the decision maker was optimizing for when making the decision and how we evaluate it. Perhaps it makes sense to make a distinction between the value function (what the decision-maker uses at the time of make a decision) and the scoring rule (which isn’t really an evaluation of the *response* but is actually an evaluation of the *value function*: does the decision-makers value function on average lead to improved outcomes).
    (2) Decision-makers face uncertain events as well as ill-expressed values. Just as interfaces can help make sense of data that *we think* correlates with future outcomes or even show predictions of future outcomes based on that data, decision “support” interfaces can also focus on helping the decision-makers express, reflect, iterate on, and communicate the otherwise ill-expressed *values* that drive their decision. This can be especially important in institutional decision-making —that Andrew often writes about—where the shared costs and benefits increase the need for justification and explanation, and where externalized values can serve to enforce consistency in repeated decisions.

    • Hi Pranav,

      Canonical lit on scoring rules evaluates relative to the state that realizes, though one could use different scoring rules for incentives versus evaluation.

      >Just as interfaces can help make sense of data that *we think* correlates with future outcomes or even show predictions of future outcomes based on that data, decision “support” interfaces can also focus on helping the decision-makers express, reflect, iterate on, and communicate the otherwise ill-expressed *values* that drive their decision

      Yes, and some of the expert elicitation literature gets into this (see Tony O’Hagan’s work for example). It is funny how at least in computer science applications of statistical modeling, we almost always take the utility / loss function as given and perfectly elicited and don’t worry about the cases where we aren’t sure. I think Alex Kale, my former student who is now faculty at UChicago, is doing some user interface work on eliciting utility functions for certain scenarios.

      • Thanks for the thoughtful reply, Jessica.

        I’ve been thinking about this more from a community decision-making perspective where I’m trying to see whether (and at what fidelity) representations of individual preferences support consensus building. These are often situations for which decision-makers “think” little historical data exists and all responses seem equally legitimate. I think decision-making research in HCAI usually focuses on judgment—with how preferences are constructed as an afterthought—because: (1) they usually focus on individual decision-makers; (2) for decisions where there already is data that might be considered relevant; and (3) helping people make sense of/explore/predict using large amounts of data seems like something well-designed interfaces could help with. I also think interfaces that result from this work often support preference construction incidentally because people often solidify what they “prefer” after exploring what’s available. I like Michelle’s work on Model Sketching which at some meaningful level attempts to foreground preference construction/elicitation in developing predictive models (https://hci.stanford.edu/publications/2023/Lam_ModelSketching_CHI23.pdf). I’ll definitely check out Alex’s work to see how he approaches these questions!

Leave a Reply

Your email address will not be published. Required fields are marked *