Not quite adversarial collaboration

Someone pointed to a paper with some questionable research claims and suggested that it could be a good candidate for an adversarial replication.

It happens that I’m not particularly interested in the particular topic being studied in that paper because as far as I can tell they have zero serious theory and zero good evidence, and they’re just chasing noise. My interest in this work is primarily methodological (How is it that these researchers used the tools of scientific inquiry and the tools of statistics in order to fool themselves and others by finding patterns from noise?), sociological (How is it that these particular barren research topics get studied and publicized?), and political (Who funds this sort of work and how does it benefit?).

But to the extent that these particular researchers are still interested in the topic–I’m sure they’d disagree with my opinion that they’re trawling through noise!–I would recommend they perform a pre-registered replication of their study or one of the others in that literature. I don’t like the term “adversarial” (but see here), but I do think that multiple perspectives are a good thing, and if these researchers or others want to conduct preregistered replications, I recommend they collaborate with some outsiders who are expert in this sort of study, as one of the concerns in these sorts of studies is alternative causal pathways that could produce these results. If these researchers continue their pattern of always doing new studies and never doing preregistered replications of their own or others’ old studies, I fear that they will spend the rest of their careers chasing noise. Which makes me sad.

Let me put it another way. Consider five reasons that a study can go wrong:

1. Selection bias in reporting (forking paths, researcher degrees of freedom, p-hacking, etc)

2. Errors in the recording of data

3. Confusion about exactly how the experiment was conducted and the data were collected

4. Inappropriate statistical analysis (for example, not adjusting for biases in data collection and measurement, not accounting for correlation in space or time, etc.)

5. Failures of identification (in psychology this is sometimes called leakage; in economics or political science we’d call it endogeneity).

These are not the only ways a study can go wrong–far from it! Other potential problems include studying a null effect (hi, ESP researchers!) or studying something where the noise is orders of magnitude higher than the signal (hi, n = 3000 sex ratio researchers!) or, for that matter, outright fraud (hi, you know who you are!). But let’s just consider the 5 issues above, all of which are undeniably serious.

Preregistration addresses issues 1, 2, and 3.

Preregistration doesn’t address issue 4, but you can address issue 4 using fake-data simulation–as long as your simulation process is realistic enough.

Preregistration doesn’t address issue 5 at all. But that’s where the collaboration with expert outsiders comes in. Again, I don’t see any advantage to framing this as “adversarial”; it’s just good to have other eyes on the project, attached to people who aren’t already committed to the result.

There’s one other thing about this idea of a collaborative preregistered replication. If, as I strongly suspect, there’s no there there, then this hypothetical replication would yield results consistent with no effect. So that’s one reason why it could make sense to not do such a study. As long as you don’t try to check in this way, you can keep the ball in the air indefinitely. By saying this, I’m not suggesting that these researchers are insincere in their beliefs. It’s just easier to defend an already published study than to risk it all with the replication. Remember what happened to Motyl et al.!

12 thoughts on “Not quite adversarial collaboration

    • Wow! That is one _FUN_ article.

      I had heard that K+T’s stuff was, if not discredited, at least argued with. This does the discredited thing beautifully.

      As a bit of an AC/DC (humanities/techy stuff) bloke, I _LOVED_ this:

      “The term and also has several legitimate meanings in English besides that of a logical conjunction. To understand what and means, the human mind once again makes intelligent inferences from content and context.”

      Sheesh. (This is, of course, my standard rant here that humans really are smart and really do think.)

      • Your last claim surely deserves extensive research. It appears to me that around 80 million Americans have difficulties thinking (actually much more if I’m honest, since both ‘tribes’ display similarities). I agree that people think sometimes and in some circumstances but in others I don’t see much evidence. As for really being ‘smart,’ I have similar doubts. As I said, I think the claim requires a lot of research – at least before I’ll agree with it.

    • This is the first time I have read this Gigerenzer paper, thanks for the link.

      I had read some Kahneman before I saw a post on Reddit that said “Kahneman made a name for himself by asking trick questions and then acting shocked when people got tricked.” Then I went back and re-read Kahneman and decided that the criticism was warranted. Gigerenzer fully fleshes out how this happens, and utterly destroys the famous Linda question along the way.

      • Exactly. I had thought Kahneman to be much better than the nudge stuff, but it seems not.

        We AI types of the 1970s and 1980s were quite taken with the ideas of heuristics and knowlede packages (frames, scripts), so Kahneman seemed to be doing our homework for us. But it turns out not. We always thought that the heuristics and “restaurant frame” stuff would always swap over to full logical reasoning when even the slightest hint of something being off (e.g. when the thing you thought was a “restaurant” insisted that you buy each item in your order as a ticket from a vending machine (as some number of places in Japan do). “Oh, it’s sort of like a restaurant, but you have to pay in advance. I wonder what else will be different.”)

        But Kahneman would have had our restaurant goers going home hungry…

      • I’m glad I’m not the only one who thought that, though believing I was (more or less) made me wonder what kind of mistakes people might characteristically make that lead them to think Kahneman described e.g. the Linda question correctly. If you spend a certain amount of time online over the decades, and you enjoy discussing things like that, you discover that there are very many people who do (or a surprising number of trolls who claim they do, or maybe both). Wondering why they think that starts to seem more useful than spending your time telling people they’re wrong on the Internet.

        My peeve with some of the questions in his (and similar) book is that I’ve been in middle-school classrooms where we were asked to do things like continue number sequences, and we absolutely were outright encouraged to guess based on what we’d already seen (we were rarely given the continuation of the sequence, and rarely given more than one guess each). The experiments would seem to depend on the questions being given to people who’d never encountered them before, much less been trained to answer one way or another. (I suppose the counter to that is if we had really strong characters, we would have resisted that training even at age eleven. Maybe.)

        • I’ll play devil’s advocate here. Gigerenzer has always appealed to me and this article is no different. But he always seems to go out of his way to say he is right and Kahneman was wrong. I see them both a mostly right. He disputes Kahneman’s example as overly narrow but I don’t see it that way. Here are a few:

          The Tom problem as the base rate fallacy. It has always bothered me, as Gigerenzer says that there is no clearly right answer to the question of what Tom’s background is likely to be. We have conflicting forces – the base rates and the personal characteristics. So, G is right that people’s choices do not demonstrate a neglect of base rates. But they don’t reject that explanation either. I am not familiar with the replication studies G cites, but I imagine a better experiment to test this might be to ask people what information they need before answering the question. Also, I know there is other research showing the base rate fallacy, so I consider this unresolved and I find value in both G&K arguments.

          The Linda demonstration of failure to understand conjunctive probabilities: G claims that people react to a meaning of “probable” that is “believable” rather than a probability statement. But, that just shifts the question. Is the more detailed description really more “believable” or does this demonstrate that a more detailed description is judged to be more likely? Might that not be a sort of cognitive illusion?

          The + vs – framing of medical treatment results. G claims that with more information about the dangers of alternative treatments, the clinician’s framing might be conveying important information that people are reacting to – and this additional information belies the narrowness of K’s experiments. Well, maybe. But it also suggests that people react to subtle non-verbal cues, and may do so whether they are right or not or relevant or not.

          Every time I read G, I come away with the same impression. He is convincing and provides much insight. But I’m just not convinced that K was really “wrong,” and G seems to go out of his way to have to claim such. I prefer to see them as more complementary.

          I think there is a more fundamental issue – and this is one that I may well disagree with G about. It is whether people are fundamentally rational and exhibit heuristics that lead to correct conclusions. I am not sure I agree. I think it is more likely that this is true under some circumstances and not others. When it comes to reasoning about uncertainty, I’m more inclined to think people suffer from cognitive failures. There is a difference between uncertainties such as strange noises in the forest (which humans have evolved to interpret) and uncertainty in investment returns (which I think requires System 2 thinking, and for which System 1 is quite fallible).

        • It’s interesting that the article mentions Piaget, because this reminds me a bit of the criticism of Piaget.

          Researchers who followed Piaget criticized his theories about fundamental developmental stages in children as being more a function of language and comprehension than of true, innate stages of cognitive development. They often see that realization as effectively “debunking” Piaget.

          I see it differently. I think it’s hard to read Piaget and not think he really did discover some kind of fundamental “truths,” such as conservation of matter being a cognitive dividing line of developmental stages. Nonetheless, I do think there’s merit in the critiques that show that developmental timelines are more fluid and less clear-cut than in Piaget’s conceptualization.

        • I was also amused by Piaget reference. Especially since it wasn’t entirely negative.

          Piaget (and Sydney Brenner and various other random blokes) showed up at the MIT AI Lab* for talks and some were friends with the Minsky family in the early 70s. A story has it that Piaget went after Minsky’s (then young daughter) with his standard “permanence” (I think it was) testing and she replied “I’m too young, I don’t have permanence yet. Go talk to my brother.”

          *: I spent some time down the hall from the AI Lab, 1973-1976. Missed Piaget, but caught Brenner’s talk.

        • Dale Lehman,

          That makes sense to me. As I get older, I’ve started to take vehement criticism with a grain of salt. But sometimes being overly vehement seems to be what it takes to get heard. And on the other hand again, it frankly doesn’t bother me as much as it used to that, say, someone says the problem with humans is we process base rates irrationally. Maybe the way we process base rates is appropriate for the things we do with them in daily life. Maybe we’re not irrational to interpret questions phrased one way differently from “equivalent” questions phrased differently. We just have to realize that analyzing society wide data requires different tools. But it’s an important point that it’s a common error people do make.

Leave a Reply

Your email address will not be published. Required fields are marked *