Beware scientists who don’t dogfood it.

I dogfood it. The statistical methods in our books are the methods we use in teaching, consulting, and research; I use Stan to solve applied problems; etc.

But not everybody does. Sometimes maybe it doesn’t matter. If that notorious Stanford medical school professor wants to take cold showers or if the now-retired Cornell food guru wants to serve himself out of a perhaps-apocryphal bottomless soup bowl or if the disgraced primatologist wants to have pseudo-conversations with monkeys, or if the nudgelords want to rearrange the food in their kitchen, whatever. They should go for it.

The bigger problem comes when an entire field doesn’t eat its own dogfood.

This issue arose in a discussion with Megan Higgs and Pamela Reinagel a couple years ago regarding why the replication crisis didn’t seem to be as big a problem in biology (at least of the wet lab variety) than in psychology.

We came up with the following explanation:

In biology, researchers have a clear incentive to try to replicate published work, because they’re using it in their own research. That’s what’s meant by biology being a “cumulative science.” Lab biologists eat their own (and each others’) dogfood.

Certain glamour areas of cognitive and social psychology are different. For example, consider social priming of the elderly-words-and-slow-walking variety. Psychologists publish this work, but it’s not like they’re giving themselves subliminal tapes featuring the speeches of Speedy Gonzalez. In contrast, biologists are using published biology research in order to do better biology research. Biology is cumulative, not just in the sense of new research building old research, but in the sense of methods cumulating as well.

Lots of times we’re not dogfooding it because the our research is intended for others. For example, political scientists are (usually) not practicing politicians and we’re rarely applying any of our political insights to our own work.

But when researchers in a field don’t eat their own dogfood, I can see how unreplicated and unreplicable results can flourish.

The limits of dogfooding

Just to be clear, I’m not saying that scientists should only dogfood it. I use lots of methods and tools that I was not involved in developing. Nor am I proposing that scientists should dogfood everything they do. As noted above, I use the methods in my books—but I don’t use the methods in all of my research articles. Research articles are speculative: they include some ideas that ultimately become useful and some that do not. I dogfood it with many of my published research ideas but not all of them. Sometimes my colleagues and I are producing delicious dogfood that we can eat and recommend to others (as with MRP, Stan, loo, PSIS, PPC, etc.); other times we’re constructing some intermediate product that, if not directly edible, might contribute someday to the sort of dogfood that can sustain us. That’s research!

26 thoughts on “Beware scientists who don’t dogfood it.

  1. >the replication crisis didn’t seem to be as big a problem in biology (at least of the wet lab variety) than in psychology

    I think there’s a lot of evidence against this, even for the wet lab variety. I agree that the bigger a finding it is, the more it will get replicated, but replicating in general is still very rare.

    https://elifesciences.org/articles/71601 – Replication project of 158 preclinical cancer effects; most didn’t successfully replicate
    https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.1000344 – Publication bias in stroke research -> overstating efficacy
    https://pmc.ncbi.nlm.nih.gov/articles/PMC8647586/ – most animal data in biomedical research is probably not even published

    Many more examples I could give

    • C’mon, we’re talking about a result you believe in bc Wansink published it. As if that would mean it’s (even potentially) true.
      It’s true that Wansink *claimed* that people ate more soup when bottomless. But you know and I know that even Wansink didn’t need to believe the result to try to get it published

  2. the replication crisis didn’t seem to be as big a problem in biology (at least of the wet lab variety) than in psychology.

    This is contrary to all the evidence. Every attempt at a replication project shows very low rates:

    https://www.sciencedirect.com/science/article/abs/pii/S0014488611002391

    https://www.reuters.com/article/business/healthcare-pharmaceuticals/in-cancer-science-many-discoveries-dont-hold-up-idUSBRE82R12Q/

    https://elifesciences.org/articles/75830

    It is important to understand that the expected rate of replication for properly designed studies checking for directional significance is 50% if the results are randomly generated.

    The observed rate is more like half (or even a quarter) of that due to various biases. Ie, we would be less misinformed by defunding the experiments, and instead just coming up with ideas and flipping a coin.

    • “The observed rate is more like half (or even a quarter) of [50%] due to various biases. Ie, we would be less misinformed by defunding the experiments, and instead just coming up with ideas and flipping a coin.”

      That doesn’t even make sense over there in the 5th dimension where you live. If what you wrote were true, we could get a 75-88% accuracy rate by performing the experiment, calculating the p value, and then concluding the exact opposite of what the data shows.

      • I think the complicating factor is that a large percentage of papers aren’t just performing the experiment and reporting the calculated the p-value. If researcher degrees of freedom are used (only) to push p-values lower and we only publish p < 0.05, then the published p-values will be biased only in one direction. How serious is the problem probably depends on the subfield. I think defunding science experiments is a bit too extreme.

      • We wouldn’t get the truth if we inverted the results because the problem is that some trials are researching null effect, i.e., any observed effect is 0. Or that trials are so underpowered that they have high Type S error rates. Poor statistical methods in preclinical research routinely leads to underpowered trials so we the overall success rate of clinical trials is muddied.

      • For years, I have been on a much vilified crusade because I insist that the word, “data” should take a singular verb. Yet, not only does Matt Skaggs write, “and then concluding the exact opposite of what the data shows” but also Randy, in a few comments above that, writes, “most animal data in biomedical research is probably not even published.”
        Note that in some fields, in particular, geodesy, the plural of datum is datums

        https://en.wikipedia.org/wiki/Datum_reference

        • Although that same Wikipedia article has a note stating “The plural of this sense of the word datum is datums by convention, in contrast with the other senses of the word in which data usually serves as both the plural form and the mass noun counterpart.”

    • There is a lot of evidence for poor replication in medical research. The coin flipping is an oversimplification but does get at a larger underlying point, that we are often mining noise.

      I do a simulation in my biostatistics course where we imagine 1,000 labs doing research, 500 of them are investigating a true real effect (delta = 0.25) while 500 of them are investigating noise (delta = 0). If we let only the labs that found a significant and positive (i.e., in the ‘right’ direction and above 0.05 – this I think is reasonable because for effect sizes smaller than that the required future trial is massive) effect proceed to a future trial we find that only a little more than half of them will be researching a real effect (53% on average).

      And of course, we have biased estimates to use to power a future study (which despite many recommendations against is still common practice). Using the estimated sample size based on preclinical data we find that among studies looking at a real effect we would have only roughly 22% median (37% mean) power in the labs examining a real effect. The expected percent finding an effect (p-value < 0.05 here for simplicity) is roughly 38% when there is a true effect,

      I ignore here looking more closely at what percent of hypotheses are true, whether trials use a clinically important effect size instead of a pilot effect to power a trial and the inherent error in replications. Nevertheless, a little simulation can make it obvious that the vast majority of clinical trials may not replicate because of known problems in the research pipeline.

      • The coin flipping is an oversimplification
        […]
        do a simulation in my biostatistics course where we imagine 1,000 labs doing research, 500 of them are investigating a true real effect (delta = 0.25) while 500 of them are investigating noise (delta = 0)

        Your simulation is making the oversimplification. Add another random value to those deltas to account for confounds and trivial/irrelevant effects.

        Now there should be zero “false “positives”, while false negative rate depend on sample size vs variance (ie, $$$ to get large N and clean measurements). Also some of your “real” effects (the delta=25 subset) will be magnified, others cancelled out.

        But anyway these models aren’t going to convince people there is a problem. The only way is to do the independent replications.

        In the meantime they will just continue to assert there is no problem and ignore anyone who points out the obvious.

        Even after the replication projects show a problem, they will continue ignoring it as long as the funding continues. But, uncoming students and people from other fields can be made aware.

  3. >the replication crisis didn’t seem to be as big a problem in biology (at least of the wet lab variety) than in psychology.

    I think there’s a lot of evidence against this.

    Errington, 2021, eLife – Huge preclinical cancer bio replication project of 158 effects; most didn’t replicate
    Van der Naald, 2020, BMJ Open Science – Most animals’ data in biomedical research is probably not published
    Sena, 2010, PloS Biology – Publication bias in animal studies of stroke leads to inflated efficacy

    Many many more examples.

    • +1

      I appreciate this distinction. The word “psychologist” is almost as broad as the word “doctor” and probably warrants disambiguation. Clinical psychologists tend to dogfood it, and perhaps even more important they are regularly feeding their dogfood to people who they are observing closely over a long period of time, and not only in the context of clinical trials. They often end up with large N intensive observation data that others can only dream of..

  4. “In biology, researchers have a clear incentive to try to replicate published work, because they’re using it in their own research.”
    I agree, based on what I learned from being married to a lab biologist. Lab experiments typically are far too complex to describe in a methods section of a paper, which has a couple of consequences. (1) people who want to learn a new method or preparation often have to spend a month or more in a lab that uses it, which is a kind of replication, and (2) it is easy for results to be artifacts of the preparation, so people working in a field do not take novel results too seriously until they are replicated by someone else. Consider the example of a 2010 paper in Science claiming that a bacterium found in Mono Lake could survive on arsenic in place of phosphorous. This got debunked, as described in https://phys.org/news/2012-07-scientists-nasa-arsenic-life-untrue.html, which includes statement that “The original study needed to be confirmed in order to be considered a true discovery,..”

    • Check the spinal cord replication project I linked above. The original tech/post-docs went to the second lab to help do the experiments. Result: ~15% replication rate.

      No doubt the rumors of high unpublished replication rates are related to why the published rates are so much lower than expected, even by chance. Ie, they are just cherrypicking the ones that turned out “right” by assuming the other results were artifacts.

  5. It’s surprising how many major results historically were not deliberately tested experimentally. Euler’s equations for rigid body motion didn’t undergo an experimental test per se. Yet centuries later they were used without thought to put men on the moon. Simple use in the intervening years seems to have been the test.

    The faith in Newton’s laws and the law of Gravity came from the hundreds of small anomalies successfully explained/corrected in celestial mechanics in the centuries after Newton. Again, the “test” was just “use”.

  6. By “biology” you probably mean “molecular biology.” As other commenters have pointed out, biomedical science has a very low replication rate, it’s in the same mess as the social “sciences.”

    • Many of the biomed experiments are molecular bio. Is there a molecular bio replication project?

      I have very strong doubts those replication rates are better than 50%. It is standard practice to run experiments unblinded and easy to justify throwing out unwanted results.

  7. I don’t think the ‘dogfooding’ aspect is the primary problem. I do agree that a lack of ‘practicing what you preach’ in situations where this would obviously be the logically coherent thing to do is a signal there is something wrong.

    But I think the primary problem results from the absence of the need for making the truly risky predictions, in the Popperian sense, and having a clear incentive for getting them right. My hypothesis would be that in fields where this requirement is less easy to avoid – failures to replicate out-of-sample are presumably more obvious when they manifest as, e.g., airplanes crashing – we tend to see better scientific performance.

    Perhaps it’s time to reconsider social sciences’ pervasive rejection of the desirability and feasibility of developing and evaluating substantive what-will-happen-if predictive statements about its subject, i.e., about society. Difficult, variable, context-dependent, uncertain, etc.? Of course! But perhaps a better way of changing incentives and culture then to keep relying exclusively on journal pubs and research grants as primary motivator for research.

  8. Once again disparaging the social sciences, when as noted above, there is so much guilt to go around (unfortunately). Have a chat with Elizabeth Bik about rampant and ridiculous fraud in the biosciences. We all can improve. Please don’t be the Fox News of blogs….

  9. Hmm…not sure I buy it, at least from a “methods” perspective. There are clear examples of developmental psychologists building on past work – especially methods. I was once told that developmentalists don’t study children, they study the methods used to study children. Just look at work on theory of mind – not only tweaks to the methods, but serious work on what those methods are revealing. There are other examples that involve infant focused research. My thinking here is your overgeneralizing – especially within a field that is really a hub science and incredibly diverse.

    • Lurking:

      Good point. Not all research is dogfoodable. For example, the dude who proved Fermat’s last theorem didn’t dogfood it: it’s not like you can do anything with the theorem. You might as well ask how someone could dogfood a hike up Mount Everest. Sometimes we do a challenge just because it’s there.

      And, for that matter, I don’t exactly “dogfood” the results in Red State Blue State. Those findings inform my understanding of politics, but there’s not quite something there to dogfood.

      So I guess when I’m talking about dogfooding it, I’m talking about certain sorts of work that is supposed to be useful, whether that be a research tool or something like a new diet or medical treatment or psychological intervention.

    • One thing not mentioned is that psychologists are encouraged to “branch out” rather than “build on” others work. When tenure letters circulated, it’s easy to say Mr X is “the world’s leading expert in stereotype threat”, rather than “has built his own variation on cognitive dissonance.” Let’s not argue about the psychology, since I may be confounding theories.
      The incentive is to farm your own domain rather than cumulate a body of insights. This was long ago labeled the “jingle jangle problem”

Leave a Reply

Your email address will not be published. Required fields are marked *