In the real world people have goals and beliefs. In a controlled experiment, you have to endow them

This is Jessica. A couple weeks ago I posted on the lack of standardization in how people design experiments to study judgment and decision making, especially in applied areas of research like visualization, human-centered AI, privacy and security, NLP, etc. My recommendation was that researchers should be able to define the decision problems they are studying in terms of the uncertain state on which the decision or belief report in each trial is based, the action space defining the range of allowable responses, the scoring rule used to incentivize and/or evaluate the reports, and process that generates the signals (i.e., stimuli) that inform on the state. And that not being able to define these things points to limitations in our ability to interpret the results we get.

I am still thinking about this topic, and why I feel strongly that when the participant isn’t given a clear goal to aim for in responding, i.e., one that is aligned with the reward they get on the task, it is hard to interpret the results. 

It’s fair to say that when we interpret the results of experiments involving human behavior, we tend to be optimistic about how what we observe in the experiment relates to people’s behavior in the “real world.” The default assumption is that the experiment results can help us understand how people behave in some realistic setting that the experimental task is meant to proxy for. There sometimes seems to be a divide among researchers, between a) those who believe that judgment and decision tasks studied in controlled experiments can be loosely based on real world tasks without worrying about things being well-defined in the context of the experiment and b) those who think that the experiment should provide (and communicate to participants) some unambiguously defined way to distinguish “correct” or at least “better” responses, even if we can’t necessarily show that this understanding matches some standard we expect to operate the real-world. 

From what I see, there are more researchers running controlled studies in applied fields that are in the former camp, whereas the latter perspective is more standard in behavioral economics. Those in applied fields appear to think it’s ok to put people in a situation where they are presented with some choice or asked to report their beliefs about something but without spelling out to them exactly how what they report will be evaluated or how their payment for doing the experiment will be affected. And I will admit I too have run studies that use under-defined tasks in the past. 

Here are some reasons I’ve heard for using not using a well-defined task in a study:

People won’t behave differently if I do that. People will sometimes cite evidence that behavior in experiments doesn’t seem very responsive to incentive schemes, extrapolating from this that giving people clear instructions on how they should think about their goals in responding (i.e., what constitutes good versus bad judgments or decisions) will not make a difference. So it’s perceived as valid to just present some stuff (treatments) and pose some questions and compare how people respond.

The real world version of this task is not well-defined. Imagine studying how people use dashboards giving information about a public health crisis, or election forecasts. Someone might argue that there is no single common decision or outcome to be predicted in the real world when people use such information, and even if we choose some decision like ‘should I wear a mask’ there is no clear single utility function, so it’s ok not to tell participants how their responses will be evaluated in the experiment. 

Having to understand a scoring rule will confuse people. Relatedly, people worry that constructing a task where there is some best response will require explaining complicated incentives to study participants. They might get confused, which will interfere with their “natural” judgment processes in this kind of situation. 

I do not find these reasons very satisfying. The problem is how to interpret the elicited responses. Sure, it may be true that in some situations, participants in experiments will act more or less than the same when you put some display of information on X in front of them and say “make this decision based on what you know about X” and when you display the same information and ask the same thing but you also explain exactly how you will judge the quality of their decision. But – I don’t think it matters if they act the same. There is still a difference: in the latter case where you’ve defined what a good versus bad judgment or decision is, you know that the participants know (or at least that you’ve attempted to tell them) what their goal is when responding. And ideally you’ve given them a reason to try to achieve that goal (incentives). So you can interpret their responses as their attempt at fulfilling that goal given the information they had at hand. In terms of the loss you observe in responses relative to the best possible performance, you still can’t disambiguate the effect of their not understanding the instructions from their inability to perform well on the task despite understanding it. But you can safely consider the loss you observe as reflecting an inability to do that task (in the context of the experiment) properly. (Of course, if your scoring rule isn’t proper then you shouldn’t expect them to be truthful under perfect understanding of the task. But the point is that we can be fairly specific about the unknowns). 

When you ask for some judgment or decision but don’t say anything about how that’s evaluated, you are building variation in how the participants interpret the task directly into your experiment design. You can’t say what their responses mean in any sort of normative sense, because you don’t know what scoring rule they had in mind. You can’t evaluate anything. 

Again this seems rather obvious, if you’re used to formulating statistical decision problems. But I encounter examples all around me that appear at odds with this perspective. I get the impression that it’s seen as a “subjective” decision for the researcher to make in fields like visualization or human-centered AI. I’ve heard studies that define tasks in a decision theoretic sense accused of “overcomplicating things.” But then when it’s time to interpret the results, the distinction is not acknowledged, and so researchers will engage in quasi-normative interpretation of responses to tasks that were never well defined to begin with.

This problem seems to stem from a failure to acknowledge the differences between behavior in the experimental world versus in the real world: We do experiments (almost always) to learn about human behavior in settings that we think are somehow related to real world settings. And in the real world, people have goals and prior beliefs. We might not be able to perceive what utility function each individual person is using, but we can assume that behavior is goal-directed in some way or another. Savage’s axioms and the derivation of expected utility theory tell us that for behavior to be “rationalizable”, a person’s choices should be consistent with their beliefs about the state and the payoffs they expect under different outcomes.

When people are in an experiment, the analogous real world goals and beliefs for that kind of task will not generally apply. For example, people might take actions in the real world for intrinsic value – e.g., I vote because I feel like I’m not a good citizen if I don’t vote. I consult the public health stats because I want to be perceived by others as informed. But it’s hard to motivate people to take actions based on intrinsic value in an experiment, unless the experiment is designed specifically to look at social behaviors like development of norms or to study how intrinsically motivated people appear to be to engage with certain content. So your experiment needs to give them a clear goal. Otherwise, they will make up a goal, and different people may do this in different ways. And so you should expect the data you get back to be a hot mess of heterogeneity. 

To be fair, the data you collect may well be a hot mess of heterogeneity anyway, because it’s hard to get people to interpret your instructions correctly. We have to be cautious interpreting the results of human-subjects experiments because there will usually be ambiguity about the participants’ understanding of the task. But at least with a well-defined task, we can point to a single source of uncertainty about our results. We can narrow down reasons for bad performance to either real challenges people face in doing that task or lack of understanding the instructions. When the task is not well-defined, the space of possible explanations of the results is huge. 

Another way of saying this is that we can only really learn things about behavior in the artificial world of the experiment. As much as we might want to equate it with some real world setting, extrapolating from the world of the controlled experiment to the real world will always be a leap of faith. So we better understand our experimental world. 

A challenge when you operate under this understanding is how to explain to people who have a more relaxed attitude about experiments why you don’t think that their results will be informative. One possible strategy is to tell people to try to see the task in their experiment from the perspective of an agent who is purely transactional or “rational”:

Imagine your experiment through the eyes of a purely transactional agent, whose every action is motivated by what external reward they perceive to be in it for them. (There are many such people in the world actually!) When a transactional agent does an experiment, they approach each question they are asked with their own question: How do I maximize my reward in answering this? When the task is well-defined and explained, they have no trouble figuring out what to do, and proceed with doing the experiment. 

However, when the transactional human reaches a question that they can’t determine how to maximize their reward on, because they haven’t been given enough information, they shut down. This is because they are (quite reasonably) unwilling to take a guess at what they should do when it hasn’t been made clear to them. 

But imagine that our experiment requires them to keep answering questions. How should we think about the responses they provide? 

We can imagine many strategies they might use to make up a response. Maybe they try to guess what you, as the experimenter, think is the right answer. Maybe they attempt to randomize. Maybe they can’t be bothered to think at all and they call in the nearest cat or three year old to act on their behalf. 

We could probably make this exercise more precise, but the point is that if you would not be comfortable interpreting the data you get under the above conditions, then you shouldn’t be comfortable interpreting the data you get from an experiment that uses an under-defined task.

7 thoughts on “In the real world people have goals and beliefs. In a controlled experiment, you have to endow them

  1. Running experiments in economics has been popular for quite some time. Economists almost always use monetary rewards, implicitly believing that the observed behavior is “realistic” since the rewards are “real.” Of course, the experimental setup is still only an experiment, so there is no guarantee that the conditions that are being simulated in the experiment will tell you anything about the conditions in the real world. For example, you want to understand how people behave in a repeated prisoners’ dilemma so you set up an appropriate game, run it for different number of iterations with monetary payoffs. Such experiments usually work well – but it is assumed that the observed behavior translates to a real situation such as a common property fishery. I’ve always found that a leap of faith that isn’t quite convincing – but it’s not feasible to experiment on the real thing. So, I guess like any scientific experiment, replication is necessary, and it would probably require replication under somewhat varying circumstances to address the issue of whether it translates to the intended situation such as the fishery.

    At the other extreme I think of surveys. There the survey questions can match the real situation perfectly – for example, a political poll is about a real election so there is no leap of faith from the poll to the real situation. However, it is almost impossible to provide any incentive in a real poll for respondents to tell the truth. It isn’t even easy to provide incentives for them to think carefully or deeply about the questions.

    I think of these two situations as ends of a spectrum – one has good incentives for an unreal situation and the other has poor incentives for a real situation. I don’t find either satisfying, but I’m not sure that can be avoided. But it would be helpful to have some rules or guideposts for how such research is conducted.

    • Yes, surveys are a good summary of the other end of the spectrum.

      I also think of studies of low level perception in contrast to more canonical judgment and decision making experiments, where in the former well designed incentives seem less important. But its hard to draw sharp boundaries between perception and cognition. Perhaps part of why these questions can seem subjective to expeirmenters.

  2. I’ve done some work at the less applied end of decision making research (mostly about how people estimate probabilities of simple and complex events), and while I agree with some of what Jessica says here, there’s a lot that I think is wrong: or at least, heading in the wrong direction.

    The first issue is that the model being used (rational participants driven by incentives and acting to maximise their reward) is just that: a model. For some judgement/decision-making tasks that model simply doesn’t apply, because there is no reward in the task (and indeed, no way of measuring performance so that a reward could be given). For example, suppose I give participants a bunch of CVs and ask them to rank those CVs in terms of suitability for a given job, based on their own judgment and experience. This is a fairly sensible task, and it could tell me something about people’s decision-making during hiring or whatever. However, I don’t know what the “right” ranking is, and so I can’t incentivize participants to give the “right” answers (or reward them when they do). Indeed, if I ask myself what the a participant’s goal in this task are, really all I can say is “their goal is to rank the CVs in some way” (I might wonder if, instead of telling me their true opinion of these CVs, they are giving a ranking that reflects well on them socially, or a ranking that they think I would like to see, etc.. This is an empty distinction, however: if those factors are part of their decision-making, then they are part of their decision making – and that’s what I’m trying to learn about in my experiment).

    The second thing that worries me about the idea that we should endow experimental participants with incentives and goals is: what happens when we encounter participants who do not seem to be following those incentives? Such participants can be easily identified (they don’t make decisions that are consistent with the goals and incentives we’ve endowed them with) and, because they are not following instructions (not following those goals and incentives) it seems reasonable to exclude them from analysis. However, if we exclude participants who don’t give the responses we expect (those consistent with the given goals), all we are left with is participants who do give the responses we expect; and since we know those responses already , why bother with an experiment?

    • Hi Fintan,

      >The first issue is that the model being used (rational participants driven by incentives and acting to maximise their reward) is just that: a model. For some judgement/decision-making tasks that model simply doesn’t apply, because there is no reward in the task (and indeed, no way of measuring performance so that a reward could be given). For example, suppose I give participants a bunch of CVs and ask them to rank those CVs in terms of suitability for a given job, based on their own judgment and experience. This is a fairly sensible task, and it could tell me something about people’s decision-making during hiring or whatever. However, I don’t know what the “right” ranking is, and so I can’t incentivize participants to give the “right” answers (or reward them when they do). Indeed, if I ask myself what the a participant’s goal in this task are, really all I can say is “their goal is to rank the CVs in some way” (I might wonder if, instead of telling me their true opinion of these CVs, they are giving a ranking that reflects well on them socially, or a ranking that they think I would like to see, etc.. This is an empty distinction, however: if those factors are part of their decision-making, then they are part of their decision making – and that’s what I’m trying to learn about in my experiment).

      This is exactly the kind of task I would question. In the real world, when people rank CVs they have goals – for instance to hire the best person for a job. If you are observing recruiters rank CVs for some actual job, great, I would consider that descriptive research. If you are pulling them into the lab, and giving them a pile of CVs, and saying rank these, sure, you can still describe what they appear to be doing, but you’ve removed the real world goals from the equation by bringing them into the lab. So when you go to characterize the results, if they are given no prompt on how to think about the ranking task, what exactly are you describing? This is exactly the kind of study I would consider uninterpretable by design.

      • Thanks Jessica. Possibly the reason I have a problem with your argument is because a while ago I did some research on whether people are rational in their judgements of likelihood and probability. There’s a fairly popular strand in cognitive psychology that suggests that they aren’t (usually referred to as the “heuristics and biases” literature: Kahneman and Tversky etc), and a lot of the judgment and decision making literature is about exactly that question, asking whether, where, and how people behave rationally (or not). From that perspective, you can’t simply take it as given that people are following goals and responding rationally to incentives, and design your experiments around that.

        My work was intended to give evidence that people are fundamentally rational, at least when considering simple probabilities and risks (see e.g. “Surprisingly rational: Probability theory plus noise explains biases in judgment”, 2014, Psychological Review, and a bunch of subsequent papers) , but we certainly couldn’t assume it from the start; and I think that’s the case for all decision and judgement research in that area. You might say “but I’m talking about goals and incentives, not rationality”, but the concept of a rational actor underlies the whole idea of people responding to incentives to achieve their goals.

        To give an example of an experiment where I think the idea of goals and incentives just can’t apply: in one study I did in 2013 or thereabouts I asked people a bunch of questions like this:

        what’s the probability that:
        A: …..Greece will leave the EU by 2020?
        B: …. the UK will leave the EU by 2020?
        A and B: … that the UK and Greece will leave the EU by 2020?
        A or B: …the UK or Greece will leave the EU by 2020?
        A|B … the UK will leave the EU by 2020, given that Greece leaves?
        and so on (randomised, interspersed with other questions etc.). I wanted to test the coherence of the probability judgements that people gave me: whether identities like P(A)+P(B)=P(A and B)+P(A or B), and P(A | B)P(B) = P(B | A)P(A) held in individual participant’s probability estimates for these questions (it turns out they did, very reliably). The problem is, though: what are participant’s incentives and goals in this task (beyond simply giving their best guess at an answer)? I can’t really see any; and I think we’d be missing out on useful research if we argued against such experiments because they don’t endow participants with goals and incentives.

        Your post has certainly provoked some thought in me, though! Thanks for that.

        • Interesting that you present this example. I have a hard time deciding whether the particular probability questions you list tell us anything at all. It is interesting to me to see how people manipulate probabilistic information, but whether they provide coherent responses or not doesn’t tell me anything. I do like the compound probability test (Kahneman and others provide additional examples) where people frequently find P(A and B) > P(A),and I have used that as an example of the kinds of problems people have with understanding probabilities. But now, in the context of Jessica’s post, I am not sure that it tells us anything useful at all. Does it really show that they have this problem when faced with probabilistic choices themselves? Or is it simply a demonstration that the richer the verbal story they are presented with, the higher the probability they assign to it? Now I’m not sure these experiments are really worthwhile at all.

  3. I wonder how much of what you’re talking about falls under the topic of “decision structuring”, as in “how do people come up with features of decisions like goals, options, desirability criteria, etc.”

    In typical decision-making studies (at least in economics, psychology, and neuroscience which I am familiar with), we impose/control/assume those features, but that often makes our models brittle in the real world. Take your decision to use a mask as an example, maybe the subject never thought about it until the experimenter brought it up. In that case, we could overestimate the likelihood of mask usage.

    I think I agree with you that we have to make some assumptions, but it’s also a statement of how little we know about how those structures are generated. I find it to be a hugely important but under-appreciated topic, even though it’s been studied by luminaries in the field from Simon to Newell…

Leave a Reply

Your email address will not be published. Required fields are marked *