“I work in a biology lab . . . My PI proposed a statistical test that I think is nonsense. . .”

This came in the email:

I work in a biology lab . . . My PI proposed a statistical test that I think is nonsense. If you have spare time I was wondering if you could look it over and tell me if my concerns are well founded.

I describe the test in an attached pdf, I wrote it up in LaTeX for readability. I wish it was shorter, I’m sorry, I did my best not to ramble.

To be honest, I am having a lot of arguments in the lab about statistical rigor, and I am wondering if you have advice about when it’s time to move on. But, this email is already long enough, there is not enough space to discuss my other concerns / vent about my frustrations.

For obv reasons I’ve anonymized the details, which is kinda too bad because the actual work looks super interesting; it has to do with modeling deformations of shapes, which is one of my favorite topics. In any case, I did not have the energy to try to understand everything that was going on, so I’ll just offer three general comments on hypothesis testing, checking the properties of a statistical method, and communicating to your supervisor:

1. The question at hand involved a permutation test. I generally don’t think permutation tests make much sense, as they are testing a null hypothesis of zero effect that is not particularly interesting, and the permutation typically corresponds to an unrealistic design; see section 3.3 of this paper.

2. A question arose about how a particular statistical method would perform: can certain effects of interest be separately estimated from available experimental data? Advice from the literature weren’t clear. My general recommendation here is that if you are skeptical of a statistical method and you do not think it should work well, one way to demonstrate or check this is to simulate fake data from the model and various alternatives and see how the procedure works. I think this should be more convincing than a theoretical argument.

3. Finally, what to do if your supervisor suggests a method that you think is nonsense? If you don’t have a great relationship to your supervisor, so that you can’t just say, “Hey, I think this method is nonsense!”, then my recommendation is to throw the burden of proof back in the other direction. If your lab publishes a paper using a particular statistical method, it’s the job of the authors of that paper to make a convincing case for why this method is appropriate for the job it is being asked to do. So when talking with your supervisor, you can ask why the method is being used, and you can express concern that you’re not sure the paper will be accepted by the journal, given various objections the reviewers might raise.

28 thoughts on ““I work in a biology lab . . . My PI proposed a statistical test that I think is nonsense. . .”

  1. I’d say just do whatever test, then (at the very least) proceed to put the data in an appendix. A thoughtful descriptive analysis (with commentary) would also be great since no one knows what process actually generated that data better than you.

    Even better if you can guess at some quantitative model to fit it, and even best if you can use the data to test the predictions of a previously published model.

    But yea, just put all that in the appendix and do whatever statistical test.

    Think of it more like students having AI do their homework then the teacher using AI to grade the homework. Its purpose is to provide an expected signal to the wider community while the parties involved do whatever they actually want to be doing.

  2. I sympathize with the student. Although I am now a supervisor, I am frequently invited to sit on thesis and defense committees, so I still feel that the supervisor gets the last word on data analysis. It’s disheartening to see student after student getting trained in NHST, removing any analysis that doesn’t turn out significant, and reporting those non-significant results as meaning there is no effect or no relationship. I try and get the student to defend their analysis, but I’m just told by the supervisor that my criticisms reflect that I have a different analytical philosophy and all that was done by the student is perfectly fine.

    Every day is a challenge to put into words why NHST, etc may be a weak, at best, way of drawing inference from data. It seems every collaborator requires a different phrasing to understand where I’m coming from.

  3. I also sympathize with the student! I agree with Andrew’s point #2, that simulating data is useful and persuasive.

    I don’t agree with #3; the response from the PI will be that they’ll deal with reviewer comments when they come. Also, odds are extremely high that the reviewers won’t care about statistical validity at all.

    More generally, the default view in much of biology is that a statistical test is appropriate if it yields p < 0.05 when you want it to, and one shouldn't think too much beyond that. If the scientific literature is filled with meaningless noise, so be it…

  4. I’m not sure what he is trying to do exactly, but in the natural and physical sciences, I am usually very skeptical of any conclusions based on elaborate statistics. You should be able to power the experiment to get data that can be evaluated with garden variety stats (or even no stats).

    This is different from social sciences, where it is rare to really be able to do experiments, especially well controlled ones, with high N. So, maybe you have to fish for inferences. But even then I Bayesianly am skeptical of needing super fancy stats.

    I would at least feel better if the fancy stats was following up on standard stats. (Like the standard stats showed an effect…and the fancy stats is brought in as the suspenders, after the belt, to show a concern is not a concern. E.g. doing a normality test after a stats test that assumes normality.)

    • Anon:

      You can be as skeptical as you want . . . but physics and pharmacology are two areas of the natural sciences where statistics are very useful. This might not make you feel better, but researchers in this area are in the business of getting as close to the truth as they can. Their business is not to make you feel better!

      • It’s true that statistics is very useful in some areas of physics, but I might agree that the stats that one usually sees are not “fancy”. At least not usually.

      • I’ve consulted in pharma CMC (on the ground management consulting and program management, not professor expert stuff). And not just once, years of it. I still remember part of my work was to get some of the people to just do basic data analysis. (Graphing data in time series, mean and standard deviations, etc. mean difference significance tests after a mfg chance) This plant wasn’t even doing that, for their annual reviews. And they had a masters stats lady on staff!

        Yes, NDAs are more sophisticated than manufacturing quality, but I still don’t think the industry is very stats savvy. And this is coming from someone who’s in the “known unknown” quadrant of stats knowledge. (Like I Dunning Kreuger know that there’s a lot I don’t know and to be careful.) I think FDA’s attitude (rightfully) is that fancier tests are useful to prove that some positive result (shown by the simpler tests) still holds if the stricter test is done. Not to rescue some clinicals where they couldn’t prove an effect with the simple tests.

        • Anon:

          I can well believe that, in pharma, simple statistical methods have more importance than complex methods. My point is only that, in some areas of pharma, complex statistical methods have their role. When it comes to dosing, it’s not necessarily true that “you should be able to power the experiment to get data that can be evaluated with garden variety stats (or even no stats).” First, you don’t have unlimited data on humans; second, some modeling can be needed to get a sense on what’s happening inside people’s bodies. I think it’s fine to say that simple methods are important; I just think you’re going too far if you dismiss “elaborate statistics” in general in the natural sciences.

          Laplace used elaborate statistics too, and he was a natural scientist!

        • You’re a good egg, Andy. Of course there will be some (epsilon greater than zero) times when sophisticated stats is justified. Maybe even not just epsilon, to be fair to you. It’s just I would also be wary of rescuing a trial (or any scientific assertion) based on fancy stats. There can be a temptation to misuse the big guns.

          Dosing is a great example. And thanks for mentioning it. Makes me think of radiation exposure. The safety limits in the industry are based on Nagasaki/Hiroshima data. And the model is linear. Most people suspect that this is overly conservative. So a millirem is not 1000th the danger of a rem, which is not 100th the danger of 100 rems. However the US regulators (NRC and NR) have taken the conservative position. But yes, I agree dosing, especially overdosing is tricky to study. If we are talking toxicology, I suspect that FDA will also take the NR/NRC attitude. (And then the trials are on animals, anyhow. So there’s still an uncertain extrapolation at play there…)

          In any case, maybe you would help them when they have their FDA NDA review. Several thousand a day, plus first class airfare and hotel and of course any other expenses. You could help them. And big $$ are considered reasonable for expert (as opposed to working) consulting assists. It’s not just the brain…but the brain plus the prestige. And they aren’t paying you for a lot of days.

          Of course if this is just some grad student or postdoc on a grant, not the next Novartis blockbuster, well…never mind!

          https://www.youtube.com/watch?v=OjYoNL4g5Vg

      • Or consider a thought experiment.

        Assume (Bayes condition one), that you were a superstar professor in number theory.

        Situation:

        A new interesting conjecture is stated, by someone else. After that, two things happen.

        A. A proof is issued. It is long as to be unable to be checked by humans (was partially computer written). But it has been computer checked to be valid.

        B. The first trillion instances were checked to satisfy the conjecture.

        Now the question is does A or B give you more confidence in the conjecture being valid. And even if your answer is A, I think you won’t have 100% confidence in it. And will be happy that the B was also done, to help you sleep better at night.

        ***

        So, I’m a little leery of results that are dependent on whamperdyne statistics (not simple statistics) in a clinicals. FDA might just say you should have done a bigger study. And no, ANOVA is not whamperdyne. But I’m assuming you are talking about some super stats thing that I’ve never heard of and that is almost never used in industry.

        P.s. first Anonymous was not me. :(

      • You may have an outdated definition of pharmacology. The aspects using quantitative models like pharmaco-kinetics/dynamics now play minor roles. Those date back to the days of biochem.

        People barely care about biochem in pharmacology now. It is all molecular bio and signal transduction. That field grew up in the NHST environment so is barely quantitative at all, it is primarily stringing together a bunch of significance tests.

        This paper is a good summary:

        Can a biologist fix a radio?–Or, what I learned while studying apoptosis
        […]
        At some point, David said, the field reaches a stage at which models, that seemed so complete, fall apart, predictions that were considered so obvious are found to be wrong, and attempts to develop wonder drugs largely fail. This stage is characterized by a sense of frustration at the complexity of the process, and by a sinking feeling that despite all that intense digging the promised cure-all may not materialize. In other words, the field hits the wall, even though the intensity of research remains unabated for a while, resulting in thousands of publications, many of which are contradictory or largely descriptive. The flood of publications is explained, in part, by the sheer amount of accumulated information (about 10,000 papers on apoptosis were published yearly over the last few years), which makes reviewers of the manuscripts as confused and overwhelmed as their authors. This stage can be summarized by the paradox that the more facts we learn the less we understand the process we study.

        https://pubmed.ncbi.nlm.nih.gov/12242150/

        It is no coincidence that sounds *exactly* like Meehl’s description of “soft psychology”.

        And in physics they have goofy ideas about stats too. Eg, a so-called “great mystery of phyiscs”: the matter-antimatter asymmetry problem.

        There is no problem. The theories predict an *average* result, but they compare to the observations, which are essentially measuring the fluctuations (due to pair-production and annihilation) around the average at an instant of time. It’s like saying there is a mystery that someone is 5 ft 10 when the average height is 5 ft 6.

        • “There is no problem. The theories predict an *average* result, but they compare to the observations, which are essentially measuring the fluctuations (due to pair-production and annihilation) around the average at an instant of time.”
          The theories can account for fluctuations too, so I’m not sure where you’re getting this.

        • The theories can account for fluctuations too, so I’m not sure where you’re getting this.

          The fluctuations are indeed in the theory. But then this is ignored, people wonder why there isn’t exactly as much matter as antimatter since they are produced at the same rate.

          It is like thinking the prior N out of a sequence of fair coinflips will always contain exactly 50% heads.

          I worked out a simple model a few years ago. And no one could provide a reason more complicated models would differ on that point:
          https://physics.stackexchange.com/questions/505662/why-is-matter-antimatter-asymmetry-surprising-if-asymmetry-can-be-generated-by

          So there is nothing surprising about observing an asymmetry, it is expected. It is just *assumed* a more complicated model would yield the same amount of matter/antimatter. Why, when it disagrees with what we should expect and observation?

          In fact the standard theories imply the universe stochastically cycles between matter- and antimatter-dominated universes (the term “matter” simply refers to the dominant species of particle). There are then “energy-only” universes followed by “big bangs” in between.

        • One thing I think could help avoid these types of issues is modifying the notation around units.

          Currently an individual measurement and the average of multiple such measurements have the same units. If it became standard practice to indicate when averaging has taken place, that may clear up a lot of problems.

          Kind of like the difference between writing out your quantitative ideas in prose vs the (at-the-time) new tech of equations.

        • “I worked out a simple model a few years ago. And no one could provide a reason more complicated models would differ on that point…”

          As the first answer says, your model is already in the literature and the reason that it is not taken as the obvious solution to the problem is there are 10 mechanisms that generate the observed asymmetry, including yours. No one knows which one, if any, is at play. Finding out which one is at play (making a verifiable prediction that distinguishes one from the others and actually seeing this verified) would constitute the beginnings of a solution.

        • Yep, the competing solutions are stuff like changing the laws of physics vs simply don’t commit a logical fallacy:

          Neither the standard model of particle physics nor the theory of general relativity provides a known explanation for why this should be so, and it is a natural assumption that the universe is neutral with all conserved charges.[3] The Big Bang should have produced equal amounts of matter and antimatter. Since this does not seem to have been the case, it is likely some physical laws must have acted differently or did not exist for matter and/or antimatter. Several competing hypotheses exist to explain the imbalance of matter and antimatter that resulted in baryogenesis. However, there is as of yet no consensus theory to explain the phenomenon, which has been described as “one of the great mysteries in physics”.[4]

          https://en.wikipedia.org/wiki/Baryon_asymmetry

          There is no problem or mystery to be explained. Its just a bunch of people commited a silly error then allowed it to grow to very embarrassing levels.

        • My background is more in materials, but with a little bit of practical nuclear power experience. I am a little turned off by the fetish for pop physics…and really not that much into non-useful physics like cosmology and string theory and the like. (Lasers and transistors and quantum hall effect are fine, but not whichness of why crap.)

          But (with no useful background, admitted), I’m not sure why people would expect equal amounts of antimatter and matter. Yes, it seems like charge is equal. Certainly in our near surroundings. And even astronomically, it seems like gravity and the remnants of the Big Bang explain motion of stars and galaxies and the like. Not electrical repulsion.

          But I’m not sure why anyone would expect antimatter to also be equal. I mean…our nearby universe is essentially all normal matter. In the few cases, we see antimatter in the normal world, it’s a tiny oddity, like pair production (electron/positron) because of gamma rays going near nuclei in reactors. And then the positrons find an electron and annihilate themselves. I guess there’s also some beta plus production in radioactive decay. But again, that’s pretty minor compared to all the normal matter that sits around. So, I just sort of think of antimatter different than electrical charge.

          I mean even if you think the Big Bang ought to have made equal amounts of matter and antimatter, then why is our nearby universe so full of normal matter? I mean if you think the left side of the universe is anti and the right side is normal, then why were charged particles also not segregated in this manner?

      • Maybe one more story about fancy methods.

        I remember a very difficult (layered complex) compound was being studied by a (very good) TEM user. He tried to get a paper published showing the structure of the material. (In TEM, very small grains can be studied, so “every sample is a single crystal”, which is useful when it is hard to grow a big enough crystal for single-crystal X ray diffraction, which is the gold standard. But then TEM is not exactly as powerful as X-ray diffraction on macroscopic crystals.

        Anyhow, he had a new method, developed for this compound with the issues of its complext modulate layered structure. And he tried to publish the structure derivation. And then the reviewers wouldn’t buy it. It was very abstract and hard to understand his method (maybe valid, maybe not, but hard to tell). And one reviewer said, if you are going to use a new method, do it on a KNOWN STRUCTURE, first.

        So, I can kinda understand the gut skepticism of fancy, fancy stats in the sciences. I mean sure, they might be mathematically valid. But you can see why someone would be leery of them in the context of practical problems and physical inferences. Especially if the people looking at it are not super statisticians themselves (and know their own limits, even to judging the fancy methods).

        Uh…but if we are talking ANOVA, then I don’t consider that fancy. I’m assuming you are talking about something I’ve never heard of. But even with ANOVA, I always tell people to watch out for degrees of freedom. More data better. Having too fancy of a regression or test on a small data set is dangerous.

    • I’ve experienced this with some of my collaborators, especially those who have been doing research since before I was born. In my experience (this may differ from Anon here) it seems to stem from their discomfort with not understanding what we are doing (e.g., they were trained to use t-tests and ANOVAs), and especially why we need to do it. I’m a big proponent of not doing statistics for the sake of statistics, but sometimes garden-variety stats work really poorly and in many fields the questions we are addressing are now so complicated that trials can’t be designed to answer them with simple stats, or no stats.

    • You’re probably thinking of some kind of fancy Frequentist stats because that’s most often what people are going to do with “elaborate statistics” these days.

      On the other hand, I’m generally skeptical of anything that doesn’t have a generative model. I want to know, “what do you think is going on here?” and “how well does that explanation work to explain the actual data?”

      For me what that means is:

      1) A dimensionless description of a process or mechanism or at least a hypothesized relationship
      2) Dynamics when appropriate.
      3) Parameters used to describe the process have meaningful interpretations
      4) There is some “discussion” that justifies why parameters should be in some hypothesized ranges described by some priors
      5) A posterior distribution that describes the uncertainty in the model
      6) Some kind of meaningful description of residuals and whether there are systematic issues in your model.

      This is “elaborate statistics” for most practitioners today, but is I think the bare minimum to be considered to be doing actual science as opposed to “scientistic” or “cargo cult” stuff.

  5. Andrew writes ‘If you don’t have a great relationship to your supervisor, so that you can’t just say, “Hey, I think this method is nonsense!”’. I think any supervisor or PI that is not at least open to hearing this kind of criticism, irrespective of quality of the relationship, has utterly failed as scientist and should not be given authority over younger scholars. Sure, I get that this is likely an all too common situation in practice and practical strategies are needed to deal with it. But I also think it’s important to denormalize this kind of immunization from rational criticism through hierarchy.

    • You are obviously right.

      I am (right now actually) very frustrated because my professor presented some findings 1.5 years ago and in my opinion it was a pretty clear ecological fallacy. I said that, but she didn’t care and presented the exact same findings again some days ago in front of 60 people. When I again objected, I was told that they did all the analyses I suggested as well, but they had no time to present it. So, she presented some 30 slides of rubbish instead of 5 well thought out slides:(

      Anyways, I also sympathize with the student but I don’t really have great advice. From my perspective, you sometimes have to bite the bullet and accept that you just can’t talk back to your supervisor. This will be frustrating and it does not make any sense, but I think sometimes there is no alternative. Also, you should check if maybe you made a mistake.

  6. In the absence of any details, it’s hard to know what exactly happened here, so it’s possible the supervisor is being dragged through the mud unfairly here. As an advisor, what I would expect from the student coming to me telling me that my proposed method is no good is a reasoned argument (with simulations, as Andrew suggests) showing me where I went wrong. Also, it would help a lot to come with an alternative approach that works demonstrably better on simulated data.

    Reminds me of my PhD student years, when I went to my advisor with linear mixed models as a better alternative to repeated measures ANOVA. At that time I lacked the knowledge needed to simulate data and it didn’t even occur to me that that was one useful way to go. In retrospect, I should have shown my advisor how the repeated measures ANOVA delivers overenthusiastic estimates compared to the by-trial data analyzed without aggregation using linear mixed models. A missed opportunity.

    PS My advisor did let me use LMMs instead of repeated measurss ANOVA.

  7. An alternative approach in the vein of “most problems are people problems not technical problems”:

    The student might have better success if they find an approach from the literature of another academic field that can be sold as “well established”. Something like Krumbein Roundness, or maybe a constant volume deformation model, or fractal dimension… since I don’t know the thing being measured, I can’t point to anything specific. “Appeal to authority” can often carry the day.

    It’s entirely possible that the PI is less statistically sophisticated, and that the PI is estimating that most reviewers will be similarly unsophisticated. I’ve known some biology PIs who take one look at LaTeX and immediately get worried that there could be a mistake or an oversight that’s not worth the risk.

    If the problem is genuinely interesting, the student could involve a collaborator to establish the analysis method as a separate paper.

    I also totally agree on simulated data. Show pictures of two shapes and say “you can tell the difference between these two shapes by looking – does the method succeed here?”.

  8. I think it’s great that the researcher felt the need to reach out to you and that you responded. Often times the challenging part of doing interdisciplinary work is the need to borrow expertise from others, all the while lacking access to experts. Such expertise is often borrowed in the form of poorly understood concepts taken from papers or textbooks that are sometimes implemented incorrectly. This post speaks broadly about the need for rigorous training in interdisciplinary programs, which I think is where we’re headed!

Leave a Reply

Your email address will not be published. Required fields are marked *