This is Jessica. In The Rational Agent Benchmark for Data Visualization, Yifan Wu, Ziyang Guo, Michalis Mamakos, Jason Hartline and I write:
Understanding how helpful a visualization is from experimental results is difficult because the observed performance is confounded with aspects of the study design, such as how useful the information that is visualized is for the task. We develop a rational agent framework for designing and interpreting visualization experiments. Our framework conceives two experiments with the same setup: one with behavioral agents (human subjects), and the other one with a hypothetical rational agent. A visualization is evaluated by comparing the expected performance of behavioral agents to that of a rational agent under different assumptions. Using recent visualization decision studies from the literature, we demonstrate how the framework can be used to pre-experimentally evaluate the experiment design by bounding the expected improvement in performance from having access to visualizations, and post-experimentally to deconfound errors of information extraction from errors of optimization, among other analyses.
I like this paper. Part of the motivation behind it was my feeling that even when we do our best to rigorously define a decision or judgment task for studying visualizations, there’s an inevitable dependence of the results on how we set up the experiment. In my lab we often put a lot of effort into making the results of experiments we run easier to interpret, like plotting model predictions back to data space to reason about magnitudes of effects, or comparing people’s performance on a task to simple baselines. But these steps don’t really resolve this dependence. And if we can’t even understand how surprising our results are in light of our own experiment design, then it seems even more futile to jump to speculating what our results imply for real world situations where people use visualizations.
We could summarize the problem in terms of various sources of unresolved ambiguity when experiment results are presented. Experimenters make many decisions in design–some of which they themselves may not even be aware they are making–which influence the range of possible effects we might see in the results. When studying information displays in particular, we might wonder about things like:
- The extent to which performance differences are likely to be driven by differences in the amount of relevant information displays convey for that task. For example, often different visualization strategies for showing distribution vary in how they summarize the data (e.g., means versus intervals vs density plots).
- How instrumental the information display is to doing well on the task – if one understood the problem but answered without looking at the visualization, how well would we expect them to do?
- To what extent participants in the study could be expected to be incentivized to use the display.
- What part of the process of responding to the task – extracting the information from the display, or figuring out what to do with it once it was extracted – led to observed losses in performance among study participants.
- And so on.
The status quo approach to writing results sections seems to be to let the reader form their own opinions on these questions. But as readers we’re often not in a good position to understand what we are learning unless we take the time to analyze the decision problem of the experiment carefully ourselves, assuming the authors have even presented it in enough detail to make that possible. Few readers are going to be willing and/or able to do this. So what we take away from the results of empirical studies on visualizations is noisy to say the least.
An alternative which we explore in this paper is to construct benchmarks using the experiment design to make the results more interpretable. First, we take the decision problem used in a visualization study and formulate it in decision theoretic terms of a data-generating model over an uncertain state drawn from some state space, an action chosen from some action space, a visualization strategy, and a scoring rule. (At least in theory, we shouldn’t have trouble picking up a paper describing an evaluative experiment and identifying these components, though in practice in fields where many experimenters aren’t thinking very explicitly about things like scoring rules at all, it might not be so easy). We then conceive a rational agent who knows the data-generating model and understands how the visualizations (signals) are generated, and compare this agent’s performance under different assumptions in pre-experimental and post-experimental analyses.
Pre-experimental analysis: One reason for analyzing the decision task pre-experimentally is to identify cases where we have designed an experiment to evaluate visualizations but we haven’t left a lot of room to observe differences between them, or we didn’t actually give participants an incentive to look at them. Oops! To define the value of information to the decision problem we look at the difference between the rational agent’s expected performance when they only have access to the prior versus when they know the prior and also see the signal (updating their beliefs and choosing the optimal action based on what they saw).
The value of information captures how much having access to the visualization is expected to improve performance on the task in payoff space. When there are multiple visualization strategies being compared, we calculate it using the maximally informative strategy. Pre-experimentally, we can look at the size of the value of information unit relative to the range of possible scores given by the scoring rule. If the expected difference in score from making the decision after looking at the visualization versus from the prior only is a small fraction of the range of possible scores on a trial, then we don’t have a lot of “room” to observe gains in performance (in the case of studying a single visualization strategy) or (more commonly) in comparing several visualization strategies.
We can also pre-experimentally compare the value of information to the baseline reward one expects to get for doing the experiment regardless of performance. Assuming we think people are motivated by payoffs (which is implied whenever we pay people for their participation), a value of information that is a small fraction of the expected baseline reward should make us question how likely participants are to put effort into the task.
Post-experimental analysis: The value of information also comes in handy post-experimentally, when we are trying to make sense of why our human participants didn’t do as well as the rational agent benchmark. We can look at what fraction of the value of information unit human participants achieve with different visualizations. We can also differentiate sources of error by calibrating the human responses. The calibrated behavioral score is the expected score of a rational agent who knows the prior but instead of updating from the joint distribution over the signal and the state, they update from the joint distribution over the behavioral responses and the state. This distribution may contain information that the agents were unable to act on. Calibrating (at least in the case of non-binary decision tasks) helps us see how much.
Specifically, calculating the difference between the calibrated score and the rational agent benchmark as a fraction of the value of information measures the extent to which participants couldn’t extract the task relevant information from the stimuli. Calculating the difference between the calibrated score and the expected score of human participants (e.g., as predicted by a model fit to the observed results) as a fraction of the value of information, measures the extent to which participants couldn’t choose the optimal action given the information they gained from the visualization.
There is an interesting complication to all of this: many behavioral experiments don’t endow participants with a prior for the decision problem, but the rational agent needs to know the prior. Technically the definitions of the losses above should allow for loss caused by not having the right prior. So I am simplifying slightly here.
To demonstrate how all this formalization can be useful in practice, we chose a couple prior award-winning visualization research papers and applied the framework. Both are papers I’m an author on – why create new methods if you can’t learn things about your own work? In both cases, we discovered things that the original papers did not account for, such as weak incentives to consult the visualization assuming you understood the task, and a better explanation for a disparity in visualization strategy rankings by performance for a belief versus a decision task. These were the first two papers we tried to apply the framework to, not cherry-picked to be easy targets. We’ve also already applied it in other experiments we’ve done, such as for benchmarking privacy budget allocation in visual analysis.
I continue to consider myself a very skeptical experimenter, since at the end of the day, decisions about whether to deploy some intervention in the world will always hinge on the (unknown) mapping between the world of your experiment and the real world context you’re trying to approximate. But I like the idea of making greater use of rational agent frameworks in visualization in that we can at least gain a better understanding of what our results mean in the context of the decision problem we are studying.
Jessica,
I read part of the paper fairly carefully and skimmed other parts. I sometimes understood a whole paragraph or two at a time, but a few things have me flummoxed and it would be fair to say that I may be missing the main point(s) of the paper. I guess my main confusion is that it’s a paper about visualization but I don’t really see how visualization comes into it!
Take the first example (I think it’s the first one), someone is trying to decide whether to put salt on their driveway. If they apply the salt and there is no freezing rain then they’ve wasted the salt. If they fail to apply the salt and there is freezing rain then they have a very slippery driveway that is dangerous or whatever, so there’s a big negative score. If they don’t apply the salt and there’s no freezing rain then everything is fine, no cost and no penalty. And if they apply the salt and there’s freezing rain then they made a good decision, congratulations and they get a cookie.
And then we have a forecast: the forecasted temperature is T_hat with an uncertainty of sigma. Given T_hat and sigma I can find the probability of freezing rain, (or maybe it’s the probability that it will freeze conditional on the fact that there is rain, I forget the exact problem), and I can calculate my expected score if I do or don’t salt the driveway. So far so good. But where does visualization come in?
There are some example visualizations: one shows just the predicted temperature, one shows the predicted temperature with a gradient display for the uncertainty, and one shows a bunch of realizations. So, OK, maybe I’m not told T_hat and sigma in a table, I have to extract them from the visualizations. From just the first visualization I get T_hat but I’d have to take a guess at sigma, I suppose I would just make that up based on my knowledge of weather forecasting but that doesn’t really make it a visualization task at all. From the second plot I could take a guess at sigma. From the third I could try to guess sigma or I suppose if there are enough realizations I could just count how many have freezing and how many don’t. But my performance on this task (or anyone’s performance) would require that they understand the math involved. Do we presuppose that? I would think that if you were to perform an experiment (in real life) to see which visualizations were most useful for improving people’s score, much of what you would be measuring is people’s mathematical understanding of the problem, not anything about the visualization. That seems like a big source of noise, requiring a much larger sample size than something more direct such as asking what fraction of days would be above the freezing point, according to the visualization. What’s the downside of doing the latter?
I’m guessing the paragraph above is nonsense or mostly nonsense because I’ve missed some very fundamental point.
Thank you for reading, Phil. I appreciate it!
The forecast example was meant to represent an example of an experiment we might find in the literature (and there are some like it). I agree, it would make more sense to elicit using a more fine-grained action space if we were actually doing the expeirment, but we used this to demonstrate. The other examples we apply the framework to later in the paper elicit finer grained responses.
In terms of it being more about mathematical understanding than visualization – yes, that is accurate to some extent. Our use of the rational agent makes this more of a visualization theory paper than an applied paper, in that we are more interested in constructing upper bounds than predicting behavioral responses from boundedly rational agents. But the beauty, at least to me, is that using this approach, we end up learning a fair amount about how to improve actual visualization experiments, in both design and interpretation.
I think I understand the idea. I’m just not clear on what is gained compared to the way the cavemen would have studied the problem (I include myself!): give people a bunch of different displays and ask them to answer questions like “which country had the fastest growth in population” or “what is the probability there will be freezing rain tomorrow” or something else that reduces it to a pure problem of interpreting the graphic, without entangling it with a calculation problem. One could offer $1 for each correct answer, if people need to be incentivized to take it seriously.
I’m definitely not proposing that! If that were a good way to do it, that’s what you would be proposing! But what’s wrong with it?
Hi Phil,
I would not think of these options as mutually exclusive. We use the binary action space forecast example to demonstrate, because we can visualizing the intuition behind a proper scoring rule in a way we can’t with a more fine-grained action space. However, what you are describing can be analyzed in this framework as a decision task without changing it. In other words, we are not advocating how researchers should design a decision task, we are advocating that they should analyze what they do to make sure its compatible with their research goals. What you gain from analyzing the problem using a rational agent framework like this is:
1) Ability to reason a priori about how big a difference you could possibly expect to see between different visualization strategies under ideal use. If the value of information to the task is small to begin with, your study is ‘dead in the water’ in a sense. You might still see large differences between visualizations in results, but it will be because one of them fails miserably, not because of a real advantage when they are used properly.
2) Related to above, ability to assess how incentivized a person who takes the scoring rule seriously would be to consult the visualizations given some prior beliefs
3) Once you collect data, ability to compare the performance you observe to the rational agent benchmark, to get a sense of how much better one might have done on this task (which is often different from 100% accuracy which is typically the default comparison)
4) Related to #3, you can look at the difference in performance you observe between different visualization strategies as a fraction of the total value of information to the task, as a way of contextualizing the results
5) By calibrating the performance you observe, and comparing calibrated behavioral responses to the rational agent benchmark (ie best attainable performance) and to the uncalibrated observed performance, you can get a sense of whether the problem with your visualization is people not knowing how to answer your question correctly, or people not being able to get the information from the visualizations.
At the end of the day, researchers can keep designing experiments the way you describe. We are just suggesting that they analyze their experiment design and their results from the standpoint of the value of the information to the task. Toward the end of the paper, we show how doing this on a few papers that have been regarded as highly rigorous shows that the original papers made some misleading claims.
OK, I think I get the point but I’m still not sure.
In the ‘freezing rain’ problem, suppose I am given the gradient display and I estimate the uncertainty in the temperature to be +/- 1 C, when it’s actually +/- 2 C. So what? If the problem of interpreting the graphic isn’t embedded in a larger decision task then there’s no way to quantify the practical effect of my error. This ties into your items 2 and 3, I think. Adding the decision problem makes the problem more realistic to the real world and provides a way of quantifying the practical difference between a good visualization and a poor one in a real-world application. That makes some sense, although it comes at the expense of adding quite a bit of complexity as opposed to just asking people questions that are directly about the graphic.
Anyway thanks for explaining.
As I wrote to you, Jessica, I really like this approach.
The pre-experimental analysis seems particularly important given that in some of the examples you present, it turned out there weren’t really strong incentives to use the visualization at all. Now in some cases, subjects may not realize that the incentives are quite weak either, which perhaps is its own can of worms.
Subjects not responding to the problem as it is put forth is definitely its own can of worms. Though as I think I’ve said in previous posts, I don’t find excuses for not using well-defined tasks of the sort ‘Well, people don’t respond differently whether I am careful about my scoring rule or I just pay a flat reward’ very persuasive.
It’s on my to-do list to write up an argument for how a well-defined decision problem is always superior when it comes to interpreting experiment results compared to an underspecified or badly defined one — even if people don’t respond that differently to one versus the other, and even if using a well-defined problem still doesn’t allow you to separate certain sources of loss due to limitations on what is observable. I think its mostly a philosophical argument, although there are also pragmatic considerations that I’m thinking through.
Yes, it is another can of worms. Though as I’ve mentioned on this blog before, I don’t find reasons for not caring about using a well-defined problem of the sort ‘Well, people seem to act the same whether I put effort into a good scoring rule or just pay a flat reward’ very satisfying.
It’s on my to-do list to write up an argument for why using a well defined decision problem is always superior to a badly defined one when it comes to interpreting experiment results — even if people don’t respond that differently to one versus the other, and even if using a well-defined problem still doesn’t allow you to separate certain sources of loss due to limitations on what is observable. I think its mostly a philosophical argument, but there are also pragmatic reasons I’m thinking about.
My impression so far is that the contributions to this discussion are mostly omitting the subjects from Bookstein’s Chapter 2: Consilience as a Rhetorical Strategy — yet, the concept of consilience is, to my reading, perhaps the most important part of Bookstein’s book. So I suggest that participants to this discussion spend more time reading Chapter 2, with attention to the concept of consilience, starting on p. 17.
I am intrigued. I read a bit that was available online of the chapter, but looks like I will need to go to the math library at Northwestern to read the non-redacted version. Can you say a little about how you see consilience relating to this discussion?
Here is how I see consilience as relating to this discussion: Quoting William Whewell from Bookstein’s p. 17: “The Consilience of Inductions takes place when an Induction, obtained form one class of facts, coincides with an Induction obtained from another different class. Thus Consilience is a test of the truth of the Theory in which it occurs.” In its full generality like this the concept may seem forbiddingly subjective, slippery, even mystical. … In Section 2.1, I recount the central thread of the continental drift narrative, focusing on the way it anticipates our consilience theme to come. Following that story, which is mainly visual, I turn to one particularly powerful recent discussion of consilience, E. O. Wilson’s widely read 1998 monograph of that title, which asserts that this notion is central to all of science. Section 2.2 reviews Wilson’s claim in the general context of biological reductionism in which he couched it, but goes on to note how a restriction of the topic to quantitative inference, as this book proposes to do, obviates any need for reductionism and tuns consilience from a philosophy for reductionism into a philosophy for numerical inference in a much broader context in a much broader context.
Section 2.3 reviews an assortment of standard arguments about rhetoric from the literature of quantitative science to show how they are more or less consistent with this reading of Wilson. (There is more, which I recommend if you can get a copy of the original from which the quotes are taken.)