Martin Forster, Marco Novelli, and Charlie Welch recently posted this preprint, which begins:
We use innovations from the frequentist and Bayesian decision-theoretic sequential experimental design literature to study whether, and when, recruitment to a pandemic-disrupted clinical trial should restart. We consider four frequentist and two Bayesian designs, two of which are new, and apply them to data from the UK’s ‘DISC’ trial, a publicly-funded trial whose recruitment was seriously disrupted by the COVID-19 pandemic. The results delivered by all six designs concur with the DISC trial’s results concerning treatment superiority. However, they do so with different levels of information, owing to different recommendations about restarting recruitment. Referencing work on the seven virtues of good statistical practice, we consider how confronting the same experimental data with a range of statistical models could assist policy-makers tasked with managing non-pandemic clinical trials during a future pandemic.
Christian Hennig pointed me to this because it makes reference to our Beyond Subjective and Objective in Statistics paper and our list of virtues:
I’m happy to hear that our framing of statistical practice in terms of these different virtues, as opposed to the unclear (to me) characterizations of methods as “subjective” or “objective,” has been helpful to these researchers.
I did not try to follow the details of the Forster et al. paper, so I will just offer some general thoughts on their endeavor, which is to compare frequentist and Bayesian sequential experiments in their applied context.
Many years ago Don Rubin pointed out that, although Bayesianism and frequentism may represent different philosophies of statistics, they are not directly comparable as methods. Bayesian statistics is a framework for producing inferences given data and assumptions; frequentist statistics is a framework for evaluating inferences given assumptions. The assumptions in the two approaches can be different; the key point is that Bayesian and frequentist ideas can go together. You can use Bayesian methods to create an inference (for example, an estimate or a hypothesis test or a confidence interval or a probabilistic prediction) and then use frequentist methods to evaluate it. Or you can use non-Bayesian methods to create a inference and then use the Bayesian framework to interpret it as an approximate Bayesian inference under some assumptions.
The pairing of Bayesian and frequentist ideas is not necessary. You can set up a model and do Bayesian inference without any further frequentist evaluation (beyond the automatic property of Bayesian inferences that they have the correct frequency properties when averaging over the assumed prior and data distributions). Conversely, you can perform frequentist evaluation of an inference that was not constructed from any Bayesian model; indeed you can sometimes construct purely frequentist inferences derived from some non-Bayesian principle such as minimax loss.
But, if you are considering various inferences, Bayesian or otherwise, it makes sense to look at their frequency properties. This is a point that Rubin made in his classic paper from 1984.
This is not automatic. Just as Bayesian inference is not a single thing but rather a framework (that is, the inference depends on your prior and data model, not on the observed data alone); similarly, frequentist evaluation involves choices, both in what frequency properties to look at (for example, unbiasedness is often taken as a desirable feature in estimation but does not make sense for probabilistic prediction, a point we make in chapter 4 of Bayesian Data Analysis; see the example on page 94 of BDA3) and in what assumptions to be made about the underlying system and the data-collection process. Bayesian methods are correct when averaging over the joint distribution of parameters and data, and frequency evaluation involves averaging as well: that’s why I say that Bayesians are frequentists.
Just as I don’t think all Bayesian inferences are good (you can have a model that makes no sense, or is inappropriate for the problem under study, or has mathematical artifacts that can be tricky to discover, as in section 3 of this paper), I also don’t think that all frequentist evaluations are appropriate. I’ve already mentioned the problem with applying the concept of unbiased estimation to prediction problems; also I’m on record as generally opposing the so-called Fisher exact test (see section 3.3 here) on the grounds that it corresponds to a data distribution–a generative model for data given parameters–that almost never applies in real life. I similarly don’t like classical multiple comparisons methods, because they are designed to guard against the generally irrelevant condition that all true effects are exactly zero.
So, yeah, doing good Bayesian inference is not always easy–you need to construct that generative model of parameters and data–and doing good frequentist evaluation is not always easy either: you need to think about what’s a reasonable model to use for the data distribution to average over, and come up with a range of plausible parameter values. But it can be done, and, indeed, if you want to compare methods, it must be done. Frequentist evaluation is the only game in town. Although it can be done approximately or implicitly, for example using cross validation or external validation of predictions without ever formally setting up a model to be averaged over.
P.S. The authors respond to the comment thread here.

This (new?) “frequentist inference doesn’t exist” stance is a bit confusing.
> Bayesian statistics is a framework for producing inferences given data and assumptions; frequentist statistics is a framework for evaluating inferences given assumptions.
Ok, so as suggested by the title “Bayesian inferences and frequentist evaluations” one may Bayesianily do “inferences” – whatever that may be.
> You can use Bayesian methods to create an inference (for example, an estimate or a hypothesis test or a confidence interval or a probabilistic prediction) and then use frequentist methods to evaluate it.
What does it mean to use Bayesian methods to create a “confidence interval”? Does it mean the kind of interval also known as “credible interval” that the practitioners of the so-called “frequentist inference” would wholly reject because it’s founded upon an error? In that case it seems that one can also use frequentist methods to calculate an alternative kind of “confidence interval”.
> indeed you can sometimes construct purely frequentist inferences derived from some non-Bayesian principle such as minimax loss
Ah. So maybe a “confidence interval” in the Wikipedia sense is a purely frequentist inference? Or does any use of the concept of probability/frequency in its construction count as derivation from some Bayesian principle?
In BDA3 under the heading “Confidence intervals” you include a discussion about “Bayesian posterior intervals” which continues as follows: “But there are some confidence intervals, derived purely from sampling-theory arguments, that differ considerably from Bayesian probability intervals.” That would suggest that confidence intervals (in the Wikipedia sense) fall in the “constructing purely frequentist inferences” camp.
Carlos:
I never said that frequentist inference doesn’t exist. From the above post: “you can sometimes construct purely frequentist inferences . . .”
In know, I read the fragment you quote. I also quoted it!
What I still don’t know is whether that “you can sometimes construct purely frequentist inferences” mention includes the “people construct frequentist confidence intervals all the time” case or not.
That first line in my comment refers to the apparent dichotomy in the title between “Bayesian inferences” and “frequentist something-elses” (and was a play on de Finetti’s “probability doesn’t exist”) but I didn’t intend it to eclipse the paragraphs that follow.
I’d maybe word it like “the Bayesian approach models inference directly whereas the frequentist approach models the evaluation of inference”. There’s still frequentist non-Bayesian inference (as you write) based on methods justified by statements regarding their evaluation (like e.g. a confidence probability). There can also be Bayesian evaluation, evaluating a method (not necessarily Bayesian using the same prior) over a population in which the parameters vary according to some distribution. In this case the interpretation of the Bayesian model would be frequentist as a model of a data generating process (I’m with you thinking that frequentist and Bayesian are not opposites; it is possible to use Bayesian methods and models together with a frequentist interpretation of probability).
The inference vs. evaluation issue isn’t prominent in the Forster et al. paper as far as I see. Bayesian and frequentist approaches are both used there for decision-making. Basically both treat the same problem there and both are taken seriously and explored in there difference.
Bayesianism demands that your route be one of a set of routes, frequentism demands that your destination be one of a set destinations. A practicioner can subscribe to one, the other, both, or neither. The tension arises because, in certain hard cases, there is no Bayesian route to the desired frequentist destination. In such cases, a practicioner must decide between whether to prioritize means or ends.
I’m not sure what do you mean by “frequentism”. It’s true that “frequentist inference” demands that your destination “behaves” (in a precise probabilistic sense) for any given value of the parameter of interest and provides the route to get there. I agree that it may better not to take that route because there is no good reason to go in that direction. Unfortunately, often people get there thinking that they arrived elsewhere.
The basic problem of statistical inference is deciding which data analytic procedure to use based on the properties of the various procedures on the menu. Bayesianism is the view that properties of the prior, likelihood and loss function under which a procedure is an approximate Bayes rule matter. Frequentism is the view that properties of the (procedure, parameter) distribution which hold for all values of the parameter matter.
For example, if I present you the interval [-3, 3] as an estimate of some parameter, a Bayesian is going to want to know what prior, likelihood, and loss would approximately reverse engineer however I actually came up with this interval, and they will assess it in view of this. By contrast, a frequentist will want to know what statistical relationship the procedure I used bears to the actual parameter value under repeated sampling, specifically what features of this relationship hold regardless of the actual parameter value.
For example, what is the worst-case (across parameter values) probability that intervals generated by this procedure trap the actual parameter value? A Bayesian need not care about this, and to the extent they do they are a frequentist.
IMHO the basic problem of statistical inference is deciding which values of a parameter cause your model to give “realistic” predictions.
Daniel:
Well put. “Realistic predictions” includes fitting the data (the likelihood) and also predictions for the population of hypothetical future cases (the prior).
> Frequentism is the view that properties of the (procedure, parameter) distribution which hold for all values of the parameter matter.
Thanks for the clarification. I asked because sometimes the term is also used in the “frequentist probability” sense instead of the “frequentist inference” sense and that may be confusing if someone takes one for the other. For example, “Bayesians are frequentists” (because they may be interested in the limiting frequency of something) may not be wrong but it doesn’t tell us much about the relation between Bayesian inference and frequentist inference. (Based on some old comments from you I think we share that view.)
Ram:
I would define frequentism as any form of evaluation of a statistical procedure assuming some probability distribution for the data. One form of frequentism involves assuming a particular value of the parameters–that’s the so-called point null hypothesis. Another form uses minimax, as in your comment. Another form averages over a distribution of the parameters–but the frequentist version of this does not require that this distribution be the same as a Bayesian prior distribution. There are also forms of frequentist analysis that hold some parameters fixed and give a distribution to others; indeed, that’s the standard way that frequentist analysis handles latent and missing data. Finally, pure Bayes is a form of frequentist analysis; it’s just not something we talk about much because the results are trivial: if you average over the joint distribution of the data, then Bayesian inferences using that model have perfect calibration.
I agree with you that a Bayesian need not care about minimax frequentist analysis; indeed, I don’t! But a Bayesian can very much care about other frequentist analyses, such as the properties of Bayesian inferences when evaluated over a sampling distribution conditional on a particular reasonable choice of parameters (not simply a parameter drawn at random from the prior), or when evaluated over alternative distributions of the parameters (not the same as the prior used in the analysis), or when evaluated over alternative data models (which is sometimes called robustness in frequentist terminology).
So, yeah, I agree with Rubin that Bayesians can and should be very interested in frequentist evaluations.
Also, as discussed in my above post, I agree with Rubin that frequentists can and should be very interested in Bayesian inference, because Bayesian methods are often a good way to come up with estimates, intervals, tests, etc., with good frequency properties.
Bayesian inference can work just fine without frequency evaluation, and frequentist methods can be derived without reference to Bayesian inference, but in both cases I think there is much to be gained by considering the other perspective too.
Daniel—To be clear, I don’t think my statement is *the* definitive statement of what statistical inference is about. It is just one way to frame things that helps illuminate the Bayesian v. frequentist contrast I have in mind. Whatever analysis we present results from applying some procedure to the data. We could always have used some other procedure. Ultimately, the reason we used the one we used is because of its properties, as compared with those of the alternatives. In my telling, Bayesians and frequentists are delineated based on which such properties they care about. And this makes clear one could care about one set, the other set, both, or neither.
Carlos—Agree. One can identify several different Bayesian v. frequentist controversies, the interpretation of probability being one, the desirable properties of statistical methods being another. My points exclusively concern the latter. On the former, I don’t think there is a truth of the matter. Sometimes probability is a useful model of epistemic dynamics, in which case the Bayesian interpretation seems relevant. Other times probability is a useful model of behavior under repeated sampling, in which case the frequentist interpretation seems relevant. What the right interpretation of probability is depends on what it is modeling in a given setting, as opposed to being an intrinsic quality of probabilities as such.
Ram, it seems to me that Bayesian probability is a strict superset of Frequentist probability. in Bayesian probability we first take what we know about a situation, and then create probability distributions based on that.
If we know for example that some data we will be fed comes from a validated random number generator process with strong cryptographic properties then we can treat it as having known frequencies and nothing else. Perfectly Bayesian.
But when we only know that there is some naturally occurring process where some outcomes are more reasonable to expect than others… say some ecology or economics or material strength experiments or whatever, then we can abandon pretending to know the frequencies and work with pure plausibility … Frequentist probability interpretations deny the legitimacy of this
Andrew—that is a helpful distinction. We can call frequentism as I have posed it classical frequentism, where the properties in question must hold uniformly over the parameter space. We can contrast this with Bayesian frequentism, where the properties in question must hold after averaging over some specified distribution over the parameter space.
Bayesian frequentism in this sense is going to coincide under regular conditions with Bayesianism as I defined it, simply because Bayes rules minimize posterior expected loss pointwise under regular conditions. But they can diverge under certain conditions, so this is a technically worthwhile distinction.
It is also worth distinguishing minimax frequentism from classical frequentism. Minimaxity is a classical property, but there are plenty of other classical properties that have been considered desirable (e.g., unbiased point estimation) which are distinct from minimaxity. So in my taxonomy, minimax frequentism is a species of classical frequentism, while Bayesian frequentism is largely not (unless we’re in the flat prior, regular model regime).
This discussion is indeed beating an undead horse (a zombie horse?) when it comes to my understanding of the issues, because even though I’ve seen this type of discussion a dozen times (roughly) I just cannot hold the frequentist answer/argument/thinking in my memory. This is my deficiency, it just somehow never sticks, but that’s why it’s good for this discussion to come around again: as has happened in the past, I can absorb the frequentist thinking enough for it to stick in my head at least for a while, and sometimes that’s handy.
So, Ram, I’m hoping you can walk me through a specific question from a frequentist point of view. I’m not looking to pick a fight, just genuinely unclear on how you would say I should think about this. Let’s take the following question: what is the probability that the next crewed mission to attempt to fly around the moon will be a success, with success defined as “the spacecraft carries the crew around the moon and returns them to earth without any of them dying.” From a frequentist perspective, should I try to find a set of events that are comparable to launching a crewed spacecraft to fly around the moon, and quantify how many have failed? What sequence should I be thinking about, that could allow me to assign a quantitative probability?
Thanks in advance.
BTW I do not intend this to be a gotcha. I have asked a question like this before, and gotten an answer that I found acceptable at the time (not an answer I necessarily _agreed_ with, but one that seemed sensible)…but I don’t remember what it was.
Phil—in brief, logistic regression.
To expand, the frequentist statistician would construct a data set where each row is an event that might have occurred, and is labeled as to whether it actually occurred or not. Along with this label, the data set would contain several columns, each capturing a scalar feature in terms of which the different events may differ, and which a subject matter expert collaborator judges as potentially relevant to predicting which events actually occur and which do not. The statistician would then fit a logistic regression (or some other binary classification model generating probabilitistic predictions), and apply it to the specific event in question. The predicted probability the model generates is the frequentist answer to your question.
A reasonable follow-up question is: one could construct many such data sets, and build many such models, and each would predict a different probability, so how do we know which one is the correct one? Part of the answer is old-fashioned model checking. Which one is best calibrated? Most discriminating? What is the out-of-sample predictive risk of each model? The other part of the answer is, which (data, model) pair render exchangeability of the test set event with the training set events most plausible, given our background knowledge?
To be clear, the sketch I gave is a simple one. If our data comes from a complex survey design, or has a complex time series or panel structure, or if there is missing data, or … then the modeling exercise is going to be more involved. And while it is critical that a frequentist actually build and employ a reasonable model for this task, this is not a distinguishing feature of frequentism. Unless you want to go the fully subjective Bayesian route, Bayesian statisticians ought to care whether their model is empirically calibrated, whether it renders exchangeability plausible, and whether it appropriately handles the various complexities of the data.
But if your concern is how a frequentist could possibly assign a probability to the event you described, I think it really is as straightforward as building an adequate probabilistic prediction model and applying it to the case of interest. “Adequate” is doing a lot of work, but it is work that the Bayesian statistician has to grapple with as well.
Ram,
Thanks for that answer. Could you tell me what might be a suitable “data set where each row is an event that might have occurred, and is labeled as to whether it actually occurred or not”? I think (?) the idea is to find data that are comparable (in some statistical sense) to the next manned flight around the moon, but I don’t know what those might be.
Even something like presidential elections, I feel like there are important ways in which they are not repeated samples from the same distribution in the same way we could say that about repeatedly rolling a die or spinning a roulette wheel. And then you go to launching a spacecraft around the moon, with people in it… sure, if it was the same basic design as was used for the Apollo program, maybe you could say there’s a sequence of repeats of “the same” event. but it’s been fifty years and the technology is different etc. etc.
If — I say if — an objection to Bayesian analysis is “subjectivity”, it seems to me that trying to construct a database for use in your hypothetical regression is at least as subjective.
But in any case thanks for your answer. I do agree that it’s a very good idea to look at historical data that try to capture some important features of the riskiness of new technological endeavors. Maybe the first flights of the Space Shuttle, the first flights of Apollo, the first manned SpaceX flights, the first Soyuz flights… maybe we also include some explanatory variables like the number of unmanned flights using the same launch technology, stuff like that. Fine. But is that “frequentist”?
Ram, I just flat out deny that this is a Frequentist model.
In order for it to be a Frequentist model there must be a PHYSICAL long run frequency and a probability sampling scheme you could in principle repeat infinitely, or at least, many many more times than the one time you did the project. You need to be able to say “if we redid the sampling, asking different experts their opinions on different projects, but somehow projects in the same ‘bag’ of projects, then over and over we do that thing, the frequency with which projects whose X value is near the X value of our real project succeed is close to the Y value we estimated with a certain sampling variation”
The fact that you can imagine such a “bag of projects” isn’t relevant to the real world. I can imagine a “bag of projects” with different properties. Whose bag is relevant is a Bayesian question because the bag is just a way to convert “plausibility” into “frequency in this bag” because the numbers work the same. That is to say, you are just measuring plausibility.
What your model is, is a Bayesian model of “under microstates of the world which have a similar measured value of X (where X is the expert assessment) and conditional on the data observed, I imagine that it is about Y% plausible that an unknown outcome of a hypothetical future event would be success, and I imagine that percentage as if it were draws from a big bag of imaginary projects.” That’s Bayesian logic.
I mean, Frequentists dress up Bayesian logic all the time, so this shouldn’t be that surprising, but I’m always struck by people assuming that because they can imagine some sequence of events that the logic they’re using is Frequentist logic. Bayesians sample from the posterior using RNGs too, but that doesn’t make Bayesian modeling Frequentist.
Consider instead a situation where you literally CAN and HAVE repeated the project multiple times… Say Monthly batches of pill bottles full of drugs, and you can show that the *real world* has stable frequencies. Then you can make statements like “provided the machinery didn’t get broken, the variation in the outcomes of percentage of underfilled pill bottles will behave like P(underfilled)”. Here the frequencies are not imagined plausibility dressed up as frequency, they are observed prior historical facts about the machine. The frequencies are physical!
People making up frequencies in some hypothetical bag are just laundering plausibility.
“what is the probability that the next crewed mission to attempt to fly around the moon will be a success, with success defined as “the spacecraft carries the crew around the moon and returns them to earth without any of them dying.”
This discussion is definitely above my pay grade, as I really don’t know what it means to use Bayesian reasoning on frequentist data. But I do know how engineers would arrive at the probability of mission failure. The only way is bottom-up.
This means first finding every way that the system can fail and exhaustively listing the modes in a Failure Modes and Effects Analysis. Once this is completed, every component failure that can cause mission failure is assessed for redundancy, and then redundancy is added where it can be. For example an O-ring that can leak and cause mission failure gets a second, identical seal behind it.
The military has vast tables of empirical component failure rate data, and these are then applied to the components used in the spacecraft, taking the redundancy and mission time into account. This is done all across the entire Failure Modes and Effects Analysis and the results are summed up using statistical methods. The final result is the summed odds that any mission critical part will fail during the mission in a way that terminates the mission, which corresponds to Phil’s “quantitative probability.”
This is some ugly stuff, and the final answer has a substantial error bar.
This can be contrasted with the approach Feynman took on the space shuttle disaster, which was to ask the most knowledgeable engineers in private to guess the odds of mission success, and then rolling those numbers up based primarily upon the guesses that suggested the highest rates of failure.
So is the former approach frequentist and the latter (roughly) Bayesian?
Matt, both approaches are Bayesian, but based on different models.
We have but one of these flights to calculate. The question of repetition of the whole flight is meaningless. A properly grounded Frequentist would say that with a Normal(1.286,0.313) random number generator, there is no probability associated with the mean, even if I didn’t tell you which number I typed into the computer program. That you don’t know it, is no reason to talk about its frequency in repetition, since it’s not repeated, it’s a particular number I typed into the code.
Similarly, there is no (frequentist) probability associated with the next flight around the moon. Either it works, or it doesn’t. It is a single event. If I think up 1000 other possible events we could ask questions about you can ask questions like “if I choose a random event from the bag what is the probability that event will succeed” but then you will be asking questions that 99.9% of the time don’t involve the moon flight of interest.
With the “build up the overall failure probability from what we know about the components” model we build bayesian knowledge of the components (sometimes well calibrated knowledge using tests, and indeed based on frequencies in those tests) to see what our knowledge of the entire system’s construction and the component failure implies about the overall failure. In the second scenario when you ask the senior engineers, you are asking them to provide plausibility values based on some general knowledge about the domain, and maybe some knowledge of the failure modes of large groups of combined components of the system…
Both are using probability to describe plausibility under a knowledge background. There isn’t a set of frequencies you can refer to.
Phil—the key is exchangeability. It may be that no two presidential elections are exactly alike, but that does not preclude productive modeling of presidential elections as conditionally iid given something or the other. There is a whole election forecasting industry (one to which our kind host has contributed extensively) made possible by this fact. What the appropriate data set and modeling involves depends on the application, and I will not pretend to sufficient expertise to advise on your examples, but the point would be to design these things in a way that renders exchangeability plausible given the relevant background knowledge. If no (data, model) pair can accomplish this to any degree of satisfaction, then I’m not sure what productive role statistics can perform in this application.
Is this frequentism? I am not interested in the terminological debate. We can call it schrequentism if it makes everyone feel better. But it’s not Bayes, it’s how much of statistics is practiced, and I would guess most of the relevant practicioners think they’re applying frequentism, for what that’s worth.
Daniel—I think this is helpful, because it clarifies that 90% of where we disagree comes down to a contentless terminology dispute. As I said to Phil, I am happy to lay down my arms and call what I’m describing something other than frequentism. Preserving that name for what I’m talking about is of less than no importance to me.
We do have a remaining disagreement though, which is whether what I am describing is just “laundered Bayes”. As I’ve acknowledged and is well-known, many classical procedures approximate, and are approximated by, Bayesian procedures. Which one is laundering which is a chicken-and-egg question that I don’t think is productive so I’m not going to pursue that.
That said, I’m curious what the laundering interpretation is of the following: in a standard potential outcomes model of a vanilla RCT, linear regression adjusted for pre-treatment predictors is a consistent estimator of the average treatment effect, regardless of how misspecified the linear model is. If I want a 95% confidence interval for the ATE (meaning worst-case coverage is 95+%), I can use one of the many model-robust standard errors in the literature together with large-sample normal theory of the ATE estimator. (Or I could use the nonparametric bootstrap.) The validity of this interval estimator does not require the linear model to be correctly specified. What Bayesian procedure is this laundering? The sort of exercise I’m talking about is a bog-standard example of what we can call (for the sake of avoiding endless terminology disputes) classical schrequentism: selecting an estimator based on properties of the (estimator, estimand) distribution under repeated sampling that hold for all values of the estimand. But I’m not sure what the Bayesian exercise is that this is a poor impression of; certainly not Bayesian linear regression, where posterior intervals do not provide correct coverage under misspecification even asymptotically.
All will become clear if you accept probability is simply the ratio of possible configurations you are interested in over the total possible configurations.
Then that both numerator and denominator are not possible to calculate except for the simplest examples. Finally that both bayesian and frequentist probabilities are approximations of that intractable true combinatoric probability.
Ram
I agree we shouldn’t argue over the meaning of words, so I will define the word “Frequentism” as “a system of inference using probability in which probability is *defined* as the frequency in which a thing occurs under hypothetically infinite repeated random sampling, and ‘random’ is defined in terms of Per Martin-Lof’s definition of high complexity sequences. In particular a random bitstream is one which is Kolmogorov incompressible.”
All of the proper applications of this system involve connecting real world events to mathematical theories relying for their mathematical validity on Kolmogorov incompressible sequences. One way we can do that is by imposing the properties of these high complexity sequences on the real world… That is by using random number generators. (Although usual RNG sequences are not incompressible, they are validated through testing to approximate the properties of those sequences well).
So, when it comes to a randomized controlled trial, now, suddenly thanks to validated software that approximates these sequences, we have the properties of those sequences superimposed on our real world, and we can rely on proofs based on the behavior of those sequences to be proofs of *real world properties* of the experimental procedures.
So, we can calculate frequency of coverage under repeated sampling and the theorems will plausibly apply to the real world events. That would include theorems about coverage of your linear regression parameter under RCT assignment etc.
Note that I suspect your example requires a discrete “treatment”. For example dosing one group with a drug of a fixed dose, and another group with a placebo. Note also that I reject the analysis of most RCTs based on the idea that they’ve given a fixed dose (measured relative to the gram). What matters *always* is a fixed dose of a drug as measured against an internal measure of the same dimension (such as the mass of the patient, or the mass of the patients blood or the mass of the patients tumor or whatever is relevant). Only dimensionless ratios of internal quantities are determinative of the outcome, not the ratio of the dose to the chosen measurement standard. And if the dose-response is nonlinear then the “linear regression adjusted for pre-treatment predictors” is *NOT* a consistent estimator of the average treatment effect regardless of the misspecification.
The deeper question is, why do we repeatedly pretend that even in the absence of RNG assignments and even with things like random digit dialing polls followed by 90% of numbers dialed being non-compliant based on choices made by the call recipient, do we continue to use proofs of frequency coverage that **do not apply** to the situation at hand, to choose how to analyze the outcomes?
The deepest question of all is, why do we think that coverage under hypothetical future repetitions that *do not occur* is a valid reason to make decisions about the world? Particularly, why do we pretend that way when Abraham Wald proved in 1947 that the only (Frequentist) admissible decision rules are Bayesian ones? https://projecteuclid.org/journals/annals-of-mathematical-statistics/volume-18/issue-4/An-Essentially-Complete-Class-of-Admissible-Decision-Functions/10.1214/aoms/1177730345.full
in summary, standard procedures in statistics fail in multiple areas of logic:
1) We fail to analyze important potentially life-saving drug trials with the basic tools of dimensionless ratios. Many people fail to even realize that they have given varying doses (measured dimensionlessly).
2) We apply mathematical theorems about high complexity sequences to real world problems where we have absolutely nothing but religious faith about their applicability.
3) Even if we want minimum frequency risk decisions, we use things like confidence intervals rather than the proven only class of admissible decision rules (Bayesian ones).
Daniel,
I can invert your question–why do we pretend that Bayesianism is anything but a laundered frequentism? Why does every sufficiently complicated Bayesian calculation use Markov Chain Monte Carlo, which is frequentist? Even a Bayesian prior is simply a regularization. Perfectly frequentist.
For every theorem that says Bayes is better in some sense, there is also a theorem that says Frequentism is better in some sense.
“1) We fail to analyze important potentially life-saving drug trials with the basic tools of dimensionless ratios. Many people fail to even realize that they have given varying doses (measured dimensionlessly).
2) We apply mathematical theorems about high complexity sequences to real world problems where we have absolutely nothing but religious faith about their applicability.
3) Even if we want minimum frequency risk decisions, we use things like confidence intervals rather than the proven only class of admissible decision rules (Bayesian ones).”
1) This has nothing to do with frequentism. Bayesian statisticians do this too. (That doesn’t make it okay, I agree about that.)
2) We guess a model and check its predictions. Is that faith? Then all science is actually faith.
3) As I said before, and as Ram said, this goes both ways. “Bayesianism is laundered frequentism” is as valid as “frequentism is laundered Bayesianism”. Even Andrew has made similar points before.
You seem to be overly ideological in your Bayesianism, far more than even respected Bayesian statisticians. Unclear why this is.
> Even a Bayesian prior is simply a regularization. Perfectly frequentist.
Whether regularization is “perfectly frequentist” depends, as much of the discussion here, on what did you mean by “frequentist”. For example, the abstract for arXiv paper 2508.03504 starts with “Classically, confidence intervals are required to have consistent coverage across all values of the parameter. However, this will inevitably break down if the underlying estimation procedure is biased.”
On the other hand if “frequentist” is synonymous with “involving probabilities” (because any probability can be interpreted as the asymptotic frequency of something) then it would be trivially true. But given that you write about “theorem[s] that says Frequentism is better in some sense” it seems that you are talking about “frequentist inference”. Which may better if you aim for something worse: why wouldn’t we prefer if possible to minimise the loss function in our situation rather than guarantee the statistical properties of the procedure in other situations?
Daniel—If you want to reserve “frequentism” for settings with literal RNG-based randomization, that is fine (peculiar, but fine). I’ve already said I’m happy to call what I’ve been describing something else. But in the vanilla RCT example I described, adjusted linear regression with robust SEs or the bootstrap can be justified by repeated-sampling properties (consistency and worst-case coverage), and that justification does not require the working linear model to be correctly specified. Your nonlinear dose-response point pertains to a different setup (it changes the estimand for one thing), and does not clarify how the workhorse classical frequentist procedure I described is stealth or sloppy Bayes.
I think it’s about time we put this discussion to bed. There is a debate here about the meaning and boundaries of frequentism. I have no dog in that fight. There is a debate here about whether all of what I have been describing is just a lame imitation of Bayes, and I think the RCT example clarifies that this is false. There is a debate about whether there is any point to doing something other than Bayes for Bayes sake. My answer is it depends on what properties you care about, which was my singular point at the outset. It doesn’t seem like we’ve made much progress on any of these disputes so it’s time to call it quits. But I appreciate the spirited exchange.
Anonymous:
When we do MCMC we use it as a computational tool to learn about Bayesian probability. it samples from a *parameter* space using a random number generator. It has nothing to do with applying frequencies to the data space. Every Bayesian agrees that if you run a RNG long enough you can determine averages from the posterior better and better. This has nothing to do with what sort of things the probabilities measure (which is not frequency of outcome in repeated DATA sampling).
My point is that soooo many people are confused about where they are doing Bayes and where they are doing Frequentist statistics. So much so that very Bayesian procedures get called Frequentist all the time.
Ram I appreciate the discussion as well. I just wanted to clarify for you because I didn’t do a good job of it. In an RCT situation where you have imposed frequency properties on the real world via the application of random number generators then you legitimately can apply repeated sampling arguments without laundering Bayes. You still may wonder why you would want to do that when there is a Frequentist argument that there always exists a Bayesian procedure that does uniformly equal to or better than the Frequentist procedure (that’s Wald’s complete class theorem).
Where Frequentist procedures launder Bayes is usually when there is not RNG determining the properties of data sampling, where the data sampling is small enough that you have no evidence of the form of a stable frequency distribution for the data (say less than several tens of thousands data points from several times and places) and they utilize maximum likelihood estimation or variants thereof (which is a Bayesian procedure in which you place a sufficiently flat prior that it doesn’t shift the maximum a-posteriori density point). I call it Bayesian, because there’s nothing but religious belief to justify the belief that the data acts “as if from a random distribution as specified in the model” and yet the procedure mimics some form of “plausibility” argument. This happens literally constantly all day long in science labs. Collect 5 mice and call them a normally distributed random variable… Collect 85 patients and call their condition IID normally distributed, or binomially distributed, or exponentially distributed or whatever. It’s not even just that we don’t know the distribution, we ALSO don’t know if the sequence of patients isn’t dependent or correlated! So frequency properties of estimators relying on IID sampling from a particular specified distributional shape simply have no reason to apply to the reality!
My point about drug trials was this… if we analyze a trial as having exactly 2 conditions, and pose the question of “what is the average difference” then we collapse a continuum of dosing down to two points (say x=0 or x=100mg), where we can simply take a difference. When all there is is a difference between two pointlike conditions, linear regression will get the right answer because two points make a line.
When appropriately analyzed, we don’t have two points, say 0mg dose and 100mg dose… What we have is a continuum, dimensionless dose varying between say 0 and 1 on some dimensionless scale (say a constant times the ratio of drug given to blood mass or some other relevant mass).
As soon as we have a continuum then when a new patient comes along, we can place them anywhere on the continuum (by varying the dose) and the relevant decision variable is where to place them on the continuum to maximize their quality of outcome. By asking the wrong question we arrive at a place where we can get a Frequentist answer. Given we’ve done an RCT we can apply the Frequentist method without needing religious belief, but we are still answering the wrong problem, which is “what dose should we give **this** patient”. Instead we answer “if we give a fixed 100mg dose, how well would patients drawn from the RCT conditions, who may be mostly very different from this patient, do on average and how often in the future if we repeated the RCT would our intervals fail to contain the actual average outcome of the trial”
Answering the wrong question with a frequency tool that isn’t laundered bayes thanks to RNGs is still answering the wrong question. So your example isn’t laundered Bayes, but it still is having a hammer and looking to make everything a nail, instead of getting a screwdriver.
Daniel—As I mentioned above, I am not committed to any one interpretation of probability in all contexts. Instead, I see probability as a mathematical model that is useful in certain situations, and what the probabilities correspond to depends on the situation. For example, if I am studying a setup where the parameter is fixed but unknown, the data is generated by a distribution belonging to a family indexed by the parameter, and I want an interval estimate for the parameter with 95% worst-case coverage, then I’m not going to have much use for distributions over the parameter space. If I instead care about average coverage wrt some distribution on the parameter, then I will (of necessity) need to utilize such a distribution. So it’s not so much that “frequentist probability” prohibits certain applications of probability theory, rather it’s that if you are a classical frequentist, you aren’t going to find many uses of probability that require giving it an epistemic interpretation.
^meant for this to be in the thread above, hazards of doing this on a smartphone.
Frequentism prohibits the use of probability in any case where an experiment isn’t repeatable in principle infinitely. For example, you have a bottle with some amount of fuel in it. You pour it into a generator, and run the generator to power some things during a power outage, the generator lasts a certain time… And now I need to have some idea what amount of fuel was originally in the bottle?
Perhaps I have measurements of the generators temperature, the ambient weather temperature, and its electrical power output through time for the entire time it was running, and I know some things about the efficiency of such engines at various temperatures, and I can work out from a model about how much input fuel is required to produce that load through time… But the fuel is gone, and the level to which the container was full is unknown, and I’m not interested in what happened to other people with other generators in other conditions. I want to know how much fuel you burned in that specific case…
Frequentism just can’t answer the question using probability.
But Bayes can answer questions about frequencies.
When it comes to “worst case coverage” we are discussing parameters that can be anywhere on the real line. If you find me a *single* real world problem where that is the case, I will pay you $100. Remember that the largest supercomputers ever built can’t even store numbers on the order of 10^80 bytes and they never will, because you’d need to store a byte per subatomic particle in the universe. 256^(10^80) is a tiny tiny number compared to the size of the real line.
Daniel:
You say, “Frequentism prohibits the use of probability in any case where an experiment isn’t repeatable in principle infinitely.” I disagree. It’s just a probability model. You could just as well say that Bayes’ theorem only applies if you are willing to make bets.
I agree that some people have defined frequentism that way, and that some frequentists have said that you can’t use probability unless there is physical randomization. But I don’t hold by that restriction, any more than I hold by the restriction that Bayes’ theorem requires betting, any more than I hold by the restriction that I can only use Euclidean geometry if I’m willing to consider lines of infinite length. These all correspond to different applications of a mathematical concept.
As to mimimax, yeah, I don’t like that either. But not all frequentism is minimax, just as not all Bayesianism is noninformative priors.
Andrew, what even is the definition of probability in the frequentism you describe then? I have always taken that definition to be “the frequency you would find in infinite replication of the data collection process” and that’s consistent with everything I’ve read about frequentist probability. How would a frequentist set up a probability model for the quantity of fuel in the canister?
I think you are being too generous.
I do agree that these things are “just models” but models have a logical structure, and they shouldn’t be applied where their logical structure doesn’t admit an answer, and as far as I can see the logical structure of the frequentist model just doesn’t apply to problems that occur all the time in science, such as “what was the mass of the star that we observed in supernova on this date?” or “is this mangled firearm the one that fired the bullet that the defendant is accused of firing?” or “did a coverup occur of a containment breach in this nuclear power plant on this day and hour in this remote location in siberia?”
Things we can’t even in principle repeat.
Daniel—again, I am not taking a position on which applications of probability theory are legitimate or not. In the problem you posed, a frequentist statistician and a Bayesian statistician would both be content to provide an interval estimate for the amount of fuel burned. Where they would differ is in the rationale for the specific interval presented. The frequentist would want the interval-generating procedure used in this case to be reliable across repeated applications. The Bayesian would want the interval to be credible under a reasonable prior and likelihood. Someone sympathetic to both frameworks would want their interval to enjoy both properties, and sometimes that is possible. Ultimately one has to decide what one is looking for in an analytic output. This decision is not aided by worrying about what sorts of probabilities are allowed under different interpretations of probability.
Ram. The data we got is unrepeatable even in principle, there can be no frequencies in the long run. The justification for a frequentist interval on the fuel burned or if the mangled gun was the one used in the murder or whatever is essentially “if this data came from a random number generator then in the long run of repetition 95% of the intervals we generate would contain the true value”.
This is the same logical justification as “if monkeys had wings then they could hover in place”. It’s based on a false premise. If you base your estimate of their caloric needs on this assumption you will get the wrong answer, or if you don’t it’s because in reality your answer is based on some other assumptions.
The fact that frequentists would simply unthinkingly apply their interval methodology and call it a day doesn’t make that a defensible position from a frequentist perspective. In most cases It will just turn out that the frequentist procedure is equivalent to some Bayesian procedure and so they won’t get a completely ridiculous answer. Just as if you assume monkeys hover in place, but then don’t use this assumption in your caloric requirement calculation and utilize only non-hovering-based information you can still get the right answer, but it’s not because you’re a “hoverist” it’s because secretly you’re a “non-hoverist”
The frequentist approach isn’t based on “if I go around with this hammer and hit various types of nails into various types of materials 95% of the nails I hit will be driven in”, it’s based on “if I hit this particular brand of nail into this particular brand of plywood over and over then 95% of the nails I hit will be driven in”. The proofs aren’t based on sampling from problems that random scientists bring you, they’re based on repeated sampling from a particular fixed frequency probability model.
For example, if you convince yourself 95% of your nails will be driven in by banging nails into plywood… and then go forth and tell people your procedure works 95% of the time, and people repeatedly bring you granite rocks and steel I-beams because they have trouble nailing things to them and they like your 95% “guarantee”, they will be disappointed and your 95% guarantee will be far from the truth.
the granite rocks and steel i-beams of statistical problems are ones that have trends rather than stationary distribution, ones that have no mean or standard deviation (cauchy etc), ones where there are serial dependencies, ones where there are systematic biases, etc etc etc. Lots of real world problems have those behaviors. Even just “polling” in the presence of adversarial non-response for example.
Anyway, we have these debates every so often. It’s useful to beat a dead horse as Andrew says because the horse is never actually dead.
Daniel—I think we’re talking past each other, since you continue to discuss interpretations of probability while I have been talking about properties of data analytic procedures.
In your example, a frequentist would write down a model of a process which the observed events instantiate. They would then apply an interval estimator which has the desired statistical properties (e.g., worst-case coverage no less than 95%) under that model. If your complaint is that it is logically impossible for the observed events to instantiate but not exhaust some data generating process, then you and I have very different ideas about what is and is not logically possible. If your complaint is that we will never observe another instantiation of this process, then I would say broaden the model to encompass instantiations we will observe, then do the regular due diligence of the statistician by making sure the model fits the data to a reasonable degree. Not sure what your objection to this is. If this isn’t allowed under your conceptual analysis of frequentism, I would say that is a flaw in your analysis rather than a flaw in frequentism.
In a “probability is plausibility” story, it is fine to say here is what we find plausible a priori, here is what it would be plausible to observe under different scenarios, so here is what this implies in terms of a posteriori plausibility. My question is what justifies the second part, the part about what is plausible to observe under different scenarios? If the answer is deduction, that the observed data follow necessarily from the parameter, then you’re not doing statistics. If your answer is unexplained behavioral dispositions as revealed in betting games, I am disposed to bet that most statisticians wouldn’t find that a compelling basis for statistical inference. The best answer, it seems to me, is that these plausibilities should be justified in the same way as the probabilities in the frequentist modeling exercise described above. If so, all this plausibility notion buys us is an interpretation of a distribution that is of no use to a (classical) frequentist (though it’s perfectly fine that it is useful from a Bayesian point of view, as I have said repeatedly).
Ram, the properties that frequentist tests have when used on random number sequences is a mathematical fact. So I don’t dispute it.
What I dispute is that these mathematical facts are relevant to the world we live in today.
It’s always possible to imagine a world that is different from our world, and run by a random number generator, and in which frequentist statistical procedures would have long run properties in use in that world for science in that world.
I just don’t choose to live in that fictional world in my head. It is an article of faith, not of science, to say that “were we in this fictional world run by random number generators, then only 0.5% of people accused of similar crimes would be wrongly convicted and therefore ___insert any statement about our real world here___”
Until a frequentist has made the connection through observation that the model of repeated sampling under random number like conditions has some good quality correspondence to our world… they are not doing science.
Note, the same is true of Bayesians, it’s just that the justification that a Bayesian needs is to justify why certain things are more plausible than other things, rather than why certain things will happen in the future with frequency f1 while others will happen with frequency f2 in long runs. These f values are more often than not asserted without even a passing attempt to connect them to reality as “standard assumptions” found in the textbook (ie. they are a social phenomenon)
This post just helped make the differences click in my head more than anything has before.
All – Thanks Andrew for mentioning our preprint and thanks to everyone for the great discussion.
I was the statistician charged with delivering and analysing data from the DISC trial, the motivation for the preprint (https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6269878). I worked on that trial before, during and after the COVID-19 pandemic and so this topic/discussion is important to me, and I think would also be of interest to other applied statisticians who might find themselves in a similar position during a future pandemic.
The three of us have read the blog contributions. Here are three observations and one question
***The observations:
1. Daniel’s discussion of “real world relevance”: the “real world” question we study is whether sequential statistical thinking can help the UK’s major funder of health research (the NIHR) do better come the next pandemic, by adopting a more nuanced approach to restarting decisions for pandemic-disrupted clinical trials investigating non-pandemic conditions. Our answer is “probably” – perhaps not all interrupted trials need to be restarted and, if that is the case, resources that are saved could be reallocated to other areas which deliver greater value e.g. saving the lives of sick patients in intensive care units or improving the resourcing of pandemic-related trials (e.g. trials of COVID-19 therapeutics, vaccines, public health interventions etc.).
2. Christian’s point about inference vs. evaluation not being prominent in our work (instead both Bayesian and frequentist approaches are there for decision-making): correct. We focus on whether the expected benefit to the UK health care system of restarting a pandemic-interrupted clinical trial outweighs the expected cost of doing so. We believe that this context quite naturally gives rise to consideration of multiple inferential and decision-making frameworks that tackle the problem from multiple perspectives.
3. Phil and Ram’s point about the value of using historical data to inform new technological endeavours: that’s precisely what our preprint aims to do. That is, use a real example/case-study to inform the frequentist and Bayesian models that we consider (including new ones), so that these tools/ideas might help policy-makers “do better” in a future pandemic.
***And here’s the question (with a big nod to Christian and Andrew’s Virtue 3(a) (“Impartiality” – Thorough consideration of relevant and potentially competing theories and points of view) and 5 (“Awareness of multiple perspectives”)):
“If the UK NIHR can improve patient health and save more lives during the next pandemic by informing trial restart decisions using virtuous statistical thinking (such as that reviewed/proposed in our preprint) and practice (e.g. fully transparent decision-making, acknowledgement of limitations etc.), does it matter whether that statistical thinking and practice is frequentist or Bayesian?”
Thanks again!
Charlie (and Martin and Marco)
Very interesting discussion! I think the idea of using virtuous statistical thinking (be it a flavor of frequentist or bayesian) and transparent processes, is a great idea. In particular, the exploration of decision theoretic properties seems to be at hand here, with expected health and financial benefits both under the microscope. The use of decision support tools to get at the better health value for money, via Bayesian or frequentist methods or both, seems to be of practical value. Thank you for exploring this issue. This would fall under the category of what i’ve seen called ‘economics of selection’ for financial-based ranking and selection which accounts for both the total expected costs of sampling and opportunity costs of potentially incorrect selections. Harken’s back to Chernoff’s 1961 paper and sequelae, and to DeGroot’s 1970 book chapter 11.8-11.9, and more recently picked up in work on health economics-oriented clinical trial design.