This is Witold.
The Food & Drug Administration just released a guidance for industry document on “Use of Bayesian Methodology in Clinical Trials of Drug and Biological Products”:
This draft guidance provides guidance to sponsors and applicants submitting investigational new drug applications (INDs), new drug applications (NDAs), biologics licensing applications (BLAs), or supplemental applications on the appropriate use of Bayesian methods in clinical trials. Bayesian methods can be used in various ways in clinical trials. For example, Bayesian calculations can be used to govern the timing and adaptation rules for an interim analysis in an adaptive design, to inform design elements (e.g., dose selection) for subsequent clinical trials, or to support primary inference in a trial. The primary focus of this guidance is on the use of Bayesian methods to support primary inference in clinical trials intended to support the effectiveness and safety of drugs.
Find it here. It’s not a very long read, about 20 pages of pretty accessible text.
I saw several people on Twitter talking it up, with someone claiming that “FDA is Bayesian now!”. Well, apparently that level of analysis is the entry bar for punditry, so that instantly made me feel qualified.
More seriously though, I bet some of the readers will have interesting takes on this, so I thought it’s good to post it and have a discussion in comments. From my side, I will try to give a short understanding of what such a document is and a few impressions.
First of all, what is a guidance document? The FDA issues and updates dozens of such documents every year. First time you encounter one, it is a bit entertaining, because it opens up with a Big Disclaimer stating not only that this sets out FDA’s “thinking” and is not binding, but even all the way down to explaining the meaning of word “should”. Why? Well, as a gov’t agency the FDA sensibly doesn’t want to give an impression of creating a law. More broadly, however, it wants to keep their options open in what they approve and what they don’t.
But then why issue guidance at all? In some sense it is pharma companies that want to have their hands tied. Drug developers are very risk averse—and for a good reason. Given the scale of investment involved in clinical trials, high uncertainty, and timelines that can stretch into decades, producers just crave predictability. So instead of taking time to interact with the FDA to vet every trial decision, it’s good to have a set of rules.
By the way, each FDA division can have its own guidance. This new one deals with drugs and vaccines. FDA released a guidance for Bayesian stats in medical device clinical trials back in 2010. This makes sense, these trials are quite different from drug or vaccine trials and people have used Bayesian techniques there for a long time, or so I’m told—I have not read that 2010 guidance (although I can report that the actual header for the doc says “baysian statistics”, sic).
Is this related to the current US gov’t and recent changes at the FDA? That’s very unlikely. (Sorry to disappoint another Twitter commenter who concluded that “Trump admin is Bayesian”.) When I blogged about this three years ago, the guidance was already scheduled to come out in 2025. Documents like this do take a long while to create.
Onto the document. As I said, I will just share a few impressions and hope that others will have more interesting comments. (Or I may also just write a second blog post on this.) Here is what jumps out scrolling through the doc:
(1) There are some basic definitions to start with, but I don’t think this matters much. It’s interesting, however, to see what type of uses the guidance lists in the second section. That probably covers the type of trials that drug and vaccine developers want to run—or that the FDA would like to see more of.
(2) The list of use cases has six categories, five of which are roughly about “borrowing”: across clinical trial phases, across subpopulations of patients, across age groups (especially extrapolating to children), or across diseases. It’s very nice that this general idea is broken down across several practical categories and given actual real-life examples in each. Lastly, there is a slightly different category of a problem, which is dose-finding, especially in cancer drugs, where you try to balance toxicity with drug’s effect.
(3) What about “fancier” adaptive types of trials? For example, trials where you add/remove treatment/dosages arms dynamically and seamlessly move across clinical trial phases. They are already covered elsewhere and the FDA is working on a separate broader guidance on “complex innovative designs”. So I don’t think there is too much about them here, although they are covered in discussion of priors.
(4) Then there are sections on success criteria and operating characteristics, which set out how to design the trials. The first proposed approach to design is… “Calibration to Type I Error Rate”. That does not sound encouraging! Frank Harrel had a great short blog post that gave a practical example of a trial where mixing Bayesian and frequentist reasoning can lead to a mess. So at first glance this is a bit disappointing…
(5) …but right next the document talks about cases where sponsor and the FDA agree to not calibrate to Type I error. There is guidance on doing things more “Bayesianly”: for example, there is even “Success Criteria Based on Benefit-Risk Assessment or Decision-Theoretic Approaches”. “A decision-theoretic approach might include assessment of the potential negative consequences of approving an ineffective drug or of not approving an effective drug.” That’s very nice! However, at the same time it’s also highly generic—the bit I just listed is limited to a single paragraph. It’s hard for me to imagine someone just reading this and confidently concluding that they can build a trial around it.
OK, I am skimming at this point, as I think this is 1,000 words already. This will probably be a follow-up blog.
(6) Then there is a long section on priors. There are many reassuring bits, because examples seem quite specific. There is a discussion of what should be considered in construction of priors. There is a discussion of “dynamic discounting”. Pretty complex and detailed. But then in talking about borrowing, the guidance says: “Areas where informative priors have been most often proposed include pediatrics and rare diseases. Additional areas can be considered on a case-by-case basis, and FDA advises early discussion of such proposals with the Agency.” Again, if this is on “case-by-case” basis, apart from two areas where we already are Bayesian, then a guidance does not seem very helpful.
So… is the FDA Bayesian now? Well, the real question should be, does a guidance document give trial designers enough confidence to run Bayesian clinical trials in situations where there is a rationale to run a Bayesian trial? Time will tell, for now there is probably a great “the guidance is directionally correct, but is it significant?” joke in there somewhere.
This is a major step in the right direction that will substantially improve our inferential and decision-making apparatus. A lot of the progress made in this draft guidance would have been unthinkable even 15 years ago for drugs and biologics.
Despite the common misconception, the FDA’s role is to provide oversight and set rigorous standards for evidence, not to design our trials for us. With that in mind, the guidance wisely avoids an over-prescriptive approach. It is also not an easy task to generate a document that can satisfy all 46,656 possible varieties of Bayesians but this one comes pretty close. Everyone involved should be very proud of this accomplishment.
Agreed! One major barrrier to Bayesian trials has been the perception that “regulators don’t like it.” So people default to what they perceive as the safe option. This should help to overcome that.
Great comment, thank you
As an update, worth juxtaposing the FDA guidance with the also recently published EMA (the European Union FDA counterpart for medicines) concept paper on Bayesian clinical trials: https://www.ema.europa.eu/en/documents/scientific-guideline/concept-paper-development-reflection-paper-use-bayesian-methods-clinical-development_en.pdf
I posted a more extensive comment on datamethods here: https://discourse.datamethods.org/t/fda-draft-guidance-use-of-bayesian-methodology-in-clinical-trials-of-drug-and-biological-products/28598/6 but the short version is that while the FDA draft guidance treats Bayesian inference as a coherent framework, the EMA still approaches it as a method requiring special justification. The EMA’s timeline to June 2028 for the final reflection paper also significantly lags behind current FDA thinking. Nevertheless, since the EMA paper is not a full guidance, but a rationale and topic list for a future reflection paper, it is a good opportunity for the statistical community to influence the EMA early.
A good and informative post.
Agreed! Thank you
Thanks both, that’s very kind. I will try to write a longer follow-up sometime, maybe once I’ve had a chance to talk with some actual trial designers
“…does a guidance document give trial designers enough confidence to run Bayesian clinical trials in situations where there is a rationale to run a Bayesian trial?”
So… are there circumstances in which there is NOT a rationale to run a Bayesian trial? To my mind bayesian methods have advantages across the board. I can’t think of anywhere I’d prefer frequentist methods.
I think regulators can use the term Bayesian in several ways. We certainty use Bayesian machinery by default for everything, but in many cases discussion of “Bayesian” in regulatory settings primarily refers to using an informative prior. You can see this more in the draft ICH E20 guidance, which is a bit more toward the Bayes=borrowing then I would prefer.
In cases where you are borrowing, FDA lately has been reasonably more worried about the data you are borrowing than then methodology used. Section V.D.2 in the guidance discusses data in detail. So in many cases the phrase “the FDA rejected Bayes” really means “the FDA rejected borrowing because they did not have comfort in the data borrowed”, which of course is different. Repeated use of the shorter phrase can be misinterpreted among regulatory consultants.
The completely open area in the new guidance is the “non type 1 error controlled world” where you will need to agree on priors and utilities. Unlike borrowing, this is essentially completely new. The guidance of course isn’t prescriptive among when they will do this. I assume the vagueness is because everyone is still figuring it out), so I will be curious and welcoming on how this evolves in practice. In the short run they will need to maintain consistency across sponsors and time (reviewers change, and bring different priors themselves). In the long run they will need to do this at scale. It will be interesting if this stays a rare “one off” for a while or a larger part of trial designs.
So few jobs and the barrier to entry is too high it’s irrelevant to waste my time thinking about. Not that I don’t understand it, and that I don’t have the ability to fit models with specifications above (I skimmed it for 2-3 seconds). But there’s really no point in investing any time into institutions that don’t pay me any money, and keep me in broke my entire life, despite codified success. Pay me. I’m not looking for suggestions that waste my time. It’s not a hobby, like many PhD’s treat it, it’s necessary I do things that put money in my pocket.
They are still talking about “calibrating to type one error rate” etc. I don’t think this is addressing any real issues.
1) Type 1 error rate is zero. The null hypothesis of no difference between groups is always false. If you fail to detect a discrepancy its because you didnt spend enough money.
– Solution, check instead for deviations from YOUR hypothesis. You derive a mechanistic model covering everything that lead to the numbers in your data. When that fits, and then is so shown to generalize to future data, we start to trust we can tell who/when/where to use the drug. You will be using “bayesian” methods for this, no need to even mention the term.
2) Bayes rule lets us compare competing hypotheses. Rather than looking at only one model in isolation, we can (and must, to be rational) compare multiple possible explanations. Ie, when looking at the discrepancies in #1, replace arbitrary thresholds with relative fit. Once again, if doing this right no need to even mention “bayesian”, because that is the way you would do it.
So, the problems are actually much simpler than that document makes it seem. The fixes are straightforward and obvious. However, I expect these will not be adopted, because it represents such a dramatic break from past practices. It amounts to admitting we need to recheck everything because the previous methods did not work (which is true but probably too embarrassing, and legally dangerous, to admit).
Setting aside the overall Bayes-vs-frequentist debate, isn’t the problem with NHST solved by the clinical trials not using point hypotheses, but instead using one-sided tests that do not start with the assumption that the performance of the drug is acceptable? That is, start with the assumption that the drug is not acceptable, then obtain enough evidence to reject this assumption. Here I’m thinking of non-inferiority or superiority approaches.
Medical colleagues I worked with sent this to me the other day. Among things I noticed is this switch: “In Bayesian approaches where the prior distribution has been chosen to provide an accurate summary of the state of belief based on existing information before the trial begins, decision-making can be based on direct interpretation of the posterior probability distribution itself. … The choice of success criteria can then be based on a determination of whether a 1 – c chance of ineffectiveness is sufficiently small in the specific context.”
I think it should say: “The choice of success criteria can then be based on a determination of whether [the/a state of belief of] 1 – c ineffectiveness is sufficiently small in the specific context.”
Deborah,
When they say, “In Bayesian approaches where the prior distribution has been chosen to provide an accurate summary of the state of belief based on existing information before the trial begins,” I want to change to “In Bayesian approaches where the prior distribution and data model have been chosen to provide an accurate summary of the state of belief based on existing information before the trial begins.”
I hate that formulation in which the prior distribution represents “the state of belief” while the data model is just there, as if it’s been handed down from the heavens.
Also, I think the idea of characterizing an intervention as “effective” or “ineffective” is generally a a bad idea, given that medical trials are typically estimating net benefits.
Not about the FDA guidance specifically, but a more general observation on these kinds of discussion about how go from individual studies to decision making (whether that’s decisions on how to design new studies, decisions on inferences about theories, or decisions about policy), from a non-expert but end user social scientist: The lack of recognition of model/theory dependence as a major inferential problem really bugs me.
It seems to me like so much of the practical advice, of the Bayesian and frequentist variety alike, just zeros in on uncertainty due to sampling variability and, to a lesser extent, measurement error. I can’t speak for other disciplines, but in the social sciences, that kind of variability is probably less of an issue than variability due to model-dependence. So, we get a lot of energy expended on meta-analysis (including the p/z-curve thing) and prior selection. But the average, however constructed, of individual model estimates based on mistaken core assumptions, is at best a noisy approximation of the truth (assuming that modeling errors somehow cancel out), at worst moving us further away from it (assuming they don’t, which often seems more plausible). Not to mention that inference drawn from any meta-analysis will itself be based on assumptions that may be wrong, as recent discussions on this blog attest to…
Hopefully I’m not misrepresenting your arguments, but you have advocated for basing inferences on severity with which claims have been tested. I think that would be a more fruitful way forward. But we’re really missing guidance on how to do this in practice. Using a concrete, simple example adapted from your chapter II in Error and Inference: a study of n=400 may provide a more severe test than a one of n=25, but it may also not if the n=400 study suffers from fatal model violations that the n=25 study does not suffer from. I really wish methodologists/statisticians would allocate more effort towards developing useful standards for assessing severity in practice, with full recognition of problem of model-dependence. Of course, there are some nice developments there, like multiverse analysis, but it seems like the focus is elsewhere.
Anyway, apologies for this rambling rant.