Witold, Erik, and I just published a paper in JAMA commenting on the recent FDA draft guidelines for the use of Bayesian methods in clinical trials. Earlier today I posted links and my reactions to two other comments on these draft guidelines published by JAMA at the same time: Embracing Bayesian Methods in Clinical Trials: FDA’s Long-Awaited Draft Guidance, by Jack Lee, Frank Harrell, Lisa LaVange, and David Spiegelhalter, and Reflections on FDA Draft Guidance on Bayesian Methods in Trials—Protecting Scientific Integrity and Evidentiary Standards, by Scott Evans, Thomas Fleming, Holly Janes, and Lori Dodd.
I pointed all the above authors to my post, and Frank Harrell shared some reactions:
I [Frank] agree completely with what you write in your blog about the guidance document and about our perspective. I also agree with your disagreements with our other cc’d hbiocolleagues except on one point: Bayesian and frequentist inference diverge when you have sequential data looks. Here is an example where the divergence is severe: https://hbiostat.org/bayes/design
I will be writing more about Tom, Scott, Holly, and Lisa’s perspective in the coming couple of weeks. Some of the perspective is written in a way that implies that FDA clinical and biostatistical reviewers are not very savvy about priors and prior data. Having worked as an FDA employee for 8 of the past 10 years I can unequivocally say that there are only a few areas where the reviewers have the wool pulled over their eyes, and these relate to being overly accepting of traditional methods used in statistical analysis plans, for example:
• allowing a sponsor to use a standard t-test on % change from baseline on a 4-point Likert pain scale
• allowing sponsors to use linear mixed effects models with no model diagnostics, e.g., not caring that an assumed compound symmetric correlation structure is correct (getting the covariance structure wrong will distort alpha)
• allowing sponsors to use change from baseline in a parallel group design (especially when the scale has floor or ceiling effects)
• not complaining when sponsors dichotomize a perfectly legitimate continuous or ordinal variable
• not asking sponsors to verify accuracy of p-value calculations for nonlinear models or models containing random effects.
I can’t comment one way or another on the real world of the FDA, but that’s a good point he has about sequential design. I gave a whole talk about this issue a couple years ago. It’s a fascinating topic that I think needs to be better understood.
Frank adds:
This example is instructive as it studies the Bayesian operating characteristics of a least-inefficient group sequential design. The findings in a nutshell are that group sequential boundaries are so conservative that when you stop with a given conclusion you can be fairly certain of its correctness, but you stop way, way too late.