This is Jessica. Occasionally on the blog the topic of “design freedoms” has come up in passing: the idea that in addition to the many degrees of freedom one can have in deciding how exactly to analyze data from some experiment, experimenters also tend to have a lot of freedom in how they design conditions, stimuli, instructions, etc, and what they choose to measure. What seems interesting is that 1) these don’t get discussed nearly as much analysis freedoms, as far as I can tell, though the problems they can pose are conceptually similar, and 2) the ability to pilot experiments repeatedly is often seen as a good thing among experimentalists. I.e., seeing something in pilot experiments then confirming it in the final experiment means you’re not reporting on your first experience observing some effect; by the time you publish it you’re replicating your own work.
I’ve never really questioned the value of pilot experiments for identifying and trying to reduce sources of noise, such as in measurement. But I’ve also suspected for a while that given all the flexibility in design conditional on seeing results through piloting, any experienced experimenter could probably take some not-so-strong effect (maybe even non-existent) and design an experiment to register it by repeatedly tweaking aspects of the design and stimuli based on pilot results.
Essentially, you have an interactive feedback process where the primary goal is to increase signal. As we know from the researcher degrees of freedom/garden of forking paths understanding, what constitutes the right signal is likely to be defined by the experimenter’s preconceptions of what they are looking for, such as a difference between conditions in a particular direction. And similarly to motivated reasoning in analysis, I could imagine some researchers not even realizing they are gaming things in a way that limits the generalizability in their results, instead thinking they are just clarifying or de-noising the observational process they’ve designed.
One obvious checkpoint in catching design gaming would seem to be peer review – you have to explain the details of your experiment to publish on it. Presumably, if you gave people in one condition some stronger directive of how to act based on the difference you want to see, reviewers should catch that, just like they should be weighing how plausibly your setting seems to capture whatever the target scenario is. E.g., for an experiment on visualizations, if you want to say something about how people interpret election graphics or other visualizations of uncertainty, you should make the experimental stimuli and task feel like something someone might go through if they encountered the visualization on the web, and you should ask them for judgments and decisions in ways they are likely to understand, and you should turn in all the instructions you have people in each condition to show that there weren’t somehow priming them with answers. But while snapshots of experimental stimuli or even code for reproducing the experiment for a reviewer to try are often submitted with experimental papers, peer review is clearly not perfect, and reviewers might miss subtle asymmetries across conditions or other artifacts of the tweaking process. Or, authors might fail to document them.
Some examples of degrees of freedom in design:
- Operationalizing the construct: In psych style studies there are often there are many ways to measure what could be construed as evidence of the same underlying construct, e.g., political attitudes, comprehension of some information, etc.
- Suggestive instructions/prompts: If you’re trying to study some effect you think is plausible, but the first, straightforward wording of the questions for subjects isn’t getting at it, you could try to walk the line between fair question and too suggestive. More generally how you explain an experimental task to people can matter a great deal in seeing effects, but interpretations often fail to acknowledge this conditionality.
- Adding validation checks: If the results seem too noisy to see what you want, try adding some more validation questions. Multiple validation questions can give the appearance that the experimenter is simply being careful to omit people who might not be reading instructions. But you could also use it to omit participants who are using more heuristic or gist-based reasoning that would make the difference between conditions look smaller.
Then there’s the whole question of choosing stimuli, which introduces lots of generalizability issues that too often get overlooked.
- Cherrypicking stimuli: The range of stimuli you test on is up to you as an experimenter. Worse, often it’s not well defined, even when it could be, which may make it harder to diagnose when it seems unreasonable.
I often think about this issue with stimuli generated to vary along some quantitative measure, since this is what happens in visualization research. For example, imagine evaluating the effect of different ways of communicating medical risks to people. How do you decide the range of risks, and how you sample from it, to create the stimuli? Ideally it’s chosen based on domain knowledge about the range of risks people face in the world in that setting (e.g., probability of some outcome if you get some disease), and sampled proportional to the expected real world distribution for whatever the target scenario is. But the chain of reasoning I see in many experimental papers (at least in my field) is weak and underspecified, more along the lines of ‘we sampled a few points along some range we think is meaningful’. An experimenter could spend some time piloting and identify some points in stimuli space that allow them to show differences, e.g., between different ways of visualizing the risk. In light of this, whether you can show an effect or even how big the effect are not at all good indicators of whether what you found was important. There are always edge cases where something counterintuitive might happen.
The importance of ample variation in stimuli is related to this. If you want to make claims about how people react to violence in video games, but then you choose some single instance of violence of a certain type, you can’t really claim to be saying something about violence in video games at large unless you have a good argument for why variation is expected to be minimal.
These conditions set up the potential for experimental results to overfit the specific context of study, i.e., a generalizability problem. So while piloting leads to replication, as Devezer et al. have made clear, this doesn’t necessarily mean that you’ve observed something important. But it does give you opportunities to tweak the choice of stimuli to your liking.
One way to better characterize this could be to simulate it. For the classic degrees of freedom in analysis, you can draw samples from the same distribution and show how you can increase your chance of seeing significance when you have flexibility around how you remove outliers, how many observations you test on, what measures that you collected end up in your final model, etc., as demonstrated in Simmons et al 2011 paper. But when piloting is used to tweak things until the experimenter is pretty sure that the “official” (perhaps pre-registered) experiment will produce results they like, issues like sampling error leading to overestimates of effects should be less of an issue. What is being exploited is some regularity induced by the experiment design, but one that isn’t really a valid proxy for the construct of interest. One way to conceptualize might be that there’s some underlying data generating process that gives rise to some moderate correlations but they can only be seen in small ranges of some total space of possible observational processes that are reasonable for studying whatever the construct of interest is. But I’m not sure that really captures it.
There’s also the obvious question of whether there should be steps to reduce potential design freedoms, like there are steps to reduce analysis freedoms. Probably not avoiding piloting experiments. We could assume that current review processes are likely to catch any important design gaming, which seems to be the current status quo. What if experimenters were also asked to report full details of the piloting process (all the things tried, screen captures of all the iterations)? Often these are tracked internally, but I haven’t really considered before what it would mean if they were submitted as part of the paper. But I would be interested as a reviewer in seeing the progression of measurement, instructions, etc choices.
All this also makes me think of blinding. When you’re piloting, even if your design uses blinding, e.g., so that an experimenter running an in-person lab experiment doesn’t know who is in what condition, the piloting process you use to tweak the design is full knowledge, and it’s hard to imagine how it wouldn’t be. Maybe if we’re going to pre-register analyses we should also pre-register piloting processes, to commit a priori to the kinds of things one might tweak and how. But that’s not really a satisfying answer as there are potentially lots of valid ways one adjusts one’s observational process that are hard to conceive of a priori.
This freedom is essential to the first abductive reasoning step of science. The problem is that people want to treat that first step as the last step without making any otherwise surprising predictions or making sure others can get the same results.
The point of independent replication is to make sure we know all the important aspects of the experimental conditions. If the piloting process is important then of course it should be included.
Yup, there is no new knowledge without abduction (explained in this allegory https://en.wikisource.org/wiki/A_Neglected_Argument_for_the_Reality_of_God)
The challenge is it needs to be good abduction and no one knows how to define good abduction.
Jessica:
I think this recent article from Gigerenzer is relevant to the points you’re making.
Yes, thanks! Hadn’t seen this one, but had seen some overlapping comments he makes in prior work, e.g., about how in many experiments on human subjects the population one thinks they are randomly sampling from is never even defined! So thinking about replication is sort of a moot point.
Very nice description of the array of issues surrounding design. It makes me think that observational studies are somewhat underrated. Sure we have the issues of file drawers, forking paths, etc. but at least the data left relatively fewer degrees of freedom (fewer, certainly not none) for the researcher to abuse.
Hadn’t thought of that, but makes sense.
This issue is frequently discussed in experimental economics. Usually, reviewers will expect that you show robustness. E.g. if you choose one particular way to present a stimulus, you’ll also have to show you’ll get the same results for other ways of presenting the stimulus.
In my field, piloting is indispensible, because studies first answer some question of What happens, and you’re then expected to explain Why it happened, using additional treatments etc. So you need to do pilots on the What before you can design appropriate treatments for the Why.
Sometimes, you’ll also need to select which elicitations (of additional preference dimensions, psychological background, etc.) you’re going to do in the limited amount of time during which subjects can be expected to pay attention. So you can’t do the relevant elicitations that allow you to speak to every possible contingency of the What that could arise. So you simply need to do pilots to prune the tree of contingencies.
The contingencies are also very hard to predict (because predicting behavior is notoriously hard), which makes meaningful preregistration of piloting strategies infeasible.
That matches my rough view of behavioral econ experiments, where sensitivity analyses do seem to be taken seriously. So perhaps the design freedoms issue is most worrisome in the weak theory paradigm, where confirming your hypotheses amounts to showing an association, maybe in a particular direction, and the need to isolate mechanisms that explain the results has mostly fallen off the agenda.
I have a related story.
I was running economic experiments simulating markets using undergraduates as test subjects.
Initially, I was using undergraduates from my school but had to expand to another school as convincing new experimenters became increasingly difficult.
What used to be data that very cleanly supported my hypothesis became a whole new level of messy.
In my mind, I blamed the students from the other school (a top 50 university in America, so easily above average, if not necessarily elite) for not understanding/following the rules and sabotaging my experiment.
But then it slowly dawned on me that there is not much point in identifying an effect that can be observed only among a group of top school undergraduates.
It was an eye-opening realization.
Fred:
That’s a good story!
Very interesting post. One related idea is the value of “stimulus sampling”, which is often a pretty vague concept, but at least suggests not just using a single word or message to instantiate some theoretical construct.
” peer review – you have to explain the details of your experiment to publish on it” I’m not so sure.
I can remember a study in which the experimental unit was a shallow pond that was seeded with sediment and a variety of aquatic plants and other organisms. Several different doses of a potential toxicant were applied to several replicates of each dose and a variety of measurements were made monthly during the growing season. On one occasion the responsible team decided to take measurements, one of which was water temperature, in treatment order on a morning when air temperature was rising rapidly. Naturally this resulted in a highly significant effect of toxicant on water temperature, which was nonsense. Did the study director put in his/her report that measurements were taken in treatment order. I doubt it.
On another occasion a contract lab carrying out an earthworm reproduction study decided, within each replicate, to rank the worms in order of weight. The heaviest worms were put in the untreated control, the next heaviest worms were given the lowest dose of the toxicant, and so on. Now, heavier worms have more offspring so this arrangement resulted in a highly significant effect of toxicant on numbers off offspring – which was not real, but an artefact of experimental methodology. Now, contract labs are very good at animal husbandry – there is no point measuring the effect of toxicants on animals that are already suffering: they have to be in peak health. So the lab would have been well aware of the correlation between weight of worm and number of offspring, but there was a lack of understanding of the purpose of experimental design, in particular sources of variation and how to control them. Randomisation often makes experiments more difficult, so experimenters will often not do it if they think it is not important.
The cause of the problem is poor training of non specialists. We emphasise analysis over principles of design. Even the text books I used early in my early career emphasised the role of randomisation of treatments, but not one of them mentioned the importance of a randomisation schedule when measuring results.
” these don’t get discussed nearly as much analysis freedoms, as far as I can tell, though the problems they can pose are conceptually similar” No,No, No. Design freedoms are much more important. A poorly designed experiment may produce data that is worthless, whilst a well designed experiment can produce useful data, often without an analysis being required.
Related: a neat suggestion from Sanjay Srivastava two days ago:
“Here’s a trick. Take the central claim of an empirical paper and preface it with, ‘A clever experimenter can carefully orchestrate circumstances in which…’ Does it still sound interesting?”
https://twitter.com/hardsci/status/1499175663420411904
(One of the rare times reading Twitter isn’t horrible or frustrating.)
Excellent post. You might find the paper on construct validity by Esterling et al interesting: https://files.osf.io/v1/resources/2s8w5/providers/osfstorage/6011b774dd222501f35923a6?format=pdf&action=download&direct&version=13
I also think we tried to elicit whether published papers pay attention to what you are describing here in our review paper (although in a very coarse way, labelled ‘special care’). https://academic.oup.com/wbro/article/33/1/34/4951685
Maybe that’s helpful.
Jörg