The webpage is maintained by John Hunter, son of Box’s collaborator William Hunter, and I came across it because I was searching for background on the paper-helicopter example that we use in our classes to teach principles of experimental design and data analysis.
There’s a lot to say about the helicopter example and I’ll save that for another post.
Here I just want to talk about how much I enjoyed reading these thirty-year-old Box articles.
A Box Set from 1990
Many of the themes in those articles continue to resonate today. For example:
• The process of learning. Here’s Box from his 1995 article, “Total Quality: Its Origins and its Future”:
Scientific method accelerated that process in at least three ways:
1. By experience in the deduction of the logical consequences of the group of facts each of which was individually known but had not previously been brought together.
2. By the passive observation of systems already in operation and the analysis of data coming from such systems.
3. By experimentation – the deliberate staging of artificial experiences which often might ordinarily never occur.
A misconception is that discovery is a “one shot” affair. This idea dies hard. . . .
• Variation over time. Here’s Box from his 1989 article, “Must We Randomize Our Experiment?”:
We all live in a non-stationary world; a world in which external factors never stay still. Indeed the idea of stationarity – of a stable world in which, without our intervention, things stay put over time – is a purely conceptual one. The concept of stationarity is useful only as a background against which the real non-stationary world can be judge. For example, the manufacture of parts is an operation involving machines and people. But the parts of a machine are not fixed entities. They are wearing out, changing their dimensions, and losing their adjustment. The behavior of the people who run the machines is not fixed either. A single operator forgets things over time and alters what he does. When a number of operators are involved, the opportunities for change because of failures to communicate are further multiplied. Thus, if left to itself any process will drift away from its initial state. . . .
Stationarity, and hence the uniformity of everything depending on it, is an unnatural state that requires a great deal of effort to achieve. That is why good quality control takes so much effort and is so important. All of this is true, not only for manufacturing processes, but for any operation that we would like to be done consistently, such as the taking of blood pressures in a hospital or the performing of chemical analyses in a laboratory. Having found the best way to do it, we would like it to be done that way consistently, but experience shows that very careful planning, checking, recalibration and sometimes appropriate intervention, is needed to ensure that this happens.
Here an example, from Box’s 1992 article, “How to Get Lucky”:
For illustration Figure 1(a) shows a set of data designed so seek out the source of unacceptably large variability which, it was suspected, might be due to small differences in five, supposedly identical, heads on a machine. To test this idea, the engineer arranged that material from each of the five heads was sampled at roughly equal intervals of time in each of six successive eight-hour periods. . . . the same analysis strongly suggested that real differences in means occurred between the six eight-hour periods of time during which the experiment was conducted. . . .
• Workflow. Here’s Box from his 1999 article, “Statistics as a Catalyst to Learning by Scientific Method Part II-Discussion”:
Most of the principles of design originally developed for agricultural experimentation would be of great value in industry, but the most industry experimentation differed from agricultural experimentation in two major respects. These I will call immediacy and sequentially.
What I mean by immediacy is that for most of our investigations the results were available, if not within hours, then certainly within days and in rare cases, even within minutes. This was true whether the investigation was conducted in a laboratory, a pilot plant or on the full scale. Furthermore, because the experimental runs were usually made in sequence, the information obtained from each run, or small group of runs, was known and could be acted upon quickly and used to plan the next set of runs. I concluded that the chief quarrel that our experimenters had with using “statistics” was that they thought it would mean giving up the enormous advantages offered by immediacy and sequentially. Quite rightly, they were not prepared to make these sacrifices. The need was to find ways of using statistics to catalyze a process of investigation that was not static, but dynamic.
There’s lots more. It’s funny to read these things that Box wrote back then, that I and others have been saying over and over again in various informal contexts, decades later. It’s a problem with our statistical education (including my own textbooks) that these important ideas are buried.
More Box
A bunch of articles by Box, with some overlap but not complete overlap with the above collection, is at the site of the University of Wisconsin, where he worked for many years. Enjoy.
Some kinda feud is going on
John Hunter’s page also has this:
The Center for Quality and Productivity Improvement was created by George Box and Bill Hunter at the University of Wisconsin-Madison in 1985.
In the first few years reports were published by leading international experts including: W. Edwards Deming, Kaoru Ishikawa, Peter Scholtes, Brian Joiner, William Hunter and George Box. William Hunter died in 1986. Subsequently excellent reports continued to be published by George Box and others including: Gipsie Ranney, Soren Bisgaard, Ron Snee and Bill Hill.
These reports were all available on the Center’s web site. After George Box’s death the reports were removed. . . .
It is a sad situation that the Center abandonded the ideas of George Box and Bill Hunter. I take what has been done to the Center as a personal insult to their memory. . . .
When diagonoised with cancer my father dedicated his remaining time to creating this center with George to promote the ideas George and he had worked on throughout their lives: because it was that important to him to do what he could. They did great work and their work provided great benefits for long after Dad’s death with the leadership of Bill Hill and Soren Bisgaard but then it deteriorated. And when George died the last restraint was eliminated and the deterioration was complete.
Wow. I wonder what the story was. I asked someone I know who works at the University of Wisconsin and he had no idea. Box died in 2013 so it’s not so long ago; there must be some people who know what happened here.
Every experiment corresponds to samples from some fixed distribution is the big lie stats textbooks tell. In reality every experiment corresponds to a process which varies in time and sometimes quite a bit. The notion that future results we be meaningfully modeled as coming from the same frequency distribution as past results is just flagrantly wrong. However the Bayesian model of “as far as I know my information about the future is symmetric with the information I have about the present…” That can be a correct model of our information. Though even there we should usually describe our information about the future as somehow more uncertain than about the present.
The phrase ‘block what you can; randomize what you cannot.’ Has been attributed to George Box. In the 1989 note ‘Must We Randomize Our Experiment’, linked above, Box is making a strong argument in favor of the *blocking* part of the phrase to potentially mitigate the effects of living in a non stationary world but a defense of *randomization* seems lacking.
If in this example of potentially having machine parts wearing out over time, why would an experimenter trying to control for that risk leaving to chance the possibility that e.g, their design assigns treatment A to the last run in each block? Why would randomization within the block be better than a systematic assignment of variable order within each block in this example’s model for the non-stationary properties of the world?
I realize that this publication is merely a note and not a fully fleshed out research paper but are good examples where the value of ‘randomizing what you cannot’ hard to construct?
The advantage of randomization is that it’s probabilistically independent of *any* factor you are unaware of. If you do things long enough, it doesn’t matter what pattern the world imposes on your experiments that pattern will balance out against your random number generator.
However this advantage really only applies to fairly large sample sizes.
So what do you think is gained by introducing a RNG to treatment assignment over an experimenter who does not have any measurement on those potential unknown confounders or a theory for how they might matter you just having to make a choice after blocking as much as they can?
You would reject the probabilistic independence on these unknown variables as a valid model for the data generating process from the experiment in the absence of using a RNG?
I can think of a couple reasons why people might still use that RNG. (1) they think it is required in order to gain the trust of others that they did not have some extra information on experimental units that they used in treatment assignment. But if people are capable of doing this they are certainly also capable of not using an RNG when they say they did. (2) frequentists might have some philosophical reasons for requiring the RNG.
But, I think one of the problems with making too much of the importance of randomization in experimental design is that folks stop looking for things they should be blocking on and randomize too much of their design.
Yes in the absence of using an RNG I would reject the idea that there’s a frequency independence between unknown factors and the treatment assignment. Note that there’s still a Bayesian independence you could apply but you’d be wrong in some sense. I can do an analysis where I assume Bayesian probabilistic independence and I’d get a valid inference under that assumption but it would be a wrong assumption.
Its like assuming that the albedo of a car is equal to the average albedo of cars on the road and then predicting the temperature. You may get a correct temperature prediction for an average car, but if the actual car is dark black, you won’t get a correct temperature prediction for the given car.
The RNG makes the Bayesian assumption correct on average over almost every possible instance of a large sample. The fact that there are few logically possible ways things could “go wrong” (ie. The RNG correlates with some hidden factor) indicates that it is almost certainly the case that we are not seeing one of those cases.
Blocking is good but randomization is needed for unknown factors.
> So what do you think is gained by introducing a RNG to treatment assignment over an experimenter who does not have any measurement on those potential unknown confounders or a theory for how they might matter you just having to make a choice after blocking as much as they can?
Is every method of making a choice – when you have to – equally valid?
Is it ok to test procedure A in the morning and procedure B in the afternoon because you don’t have a theory of how the time of the day may matter?
Is it fine to give the active treatment to patients who live with someone, or who live close to the hospital, if that simplifies the logistics?
Carlos,
“ Is it ok to test procedure A in the morning and procedure B in the afternoon because you don’t have a theory of how the time of the day may matter?”
No.
So you’ve identified something that can be blocked on. Why randomize on that factor and risk getting a bad experiment where A ends up being assigned in the morning and B in the afternoon?
“Block what you can” is the prescription to use all the known variables whether you have a theory for how they may influence the outcome or not.
But certainly when you have a theory about a variable, e.gl machine wear, that theory should be used to construct an optimal design.
Daniel,
I really don’t understand the connection between using an RNG and your ecological fallacy example.
You say the RNG ensures independence with the unknown variables, but I don’t think you’ve argued why the choice without an RNG wouldn’t be,
It seems like you’d have to take the philosophical viewpoint that “all things are dependent” to arrive at the latter, but that then leaves very little room for your RNG itself to not be entwined with the unknown variables.
JYD, the big problem is “everything affects everything else”. Should you block on the motion of a gram of mass in a star lightyears away?
https://eighteenthelephant.com/2021/11/29/pushed-around-by-stars/
If not that, then maybe the sales of different brands of dog kibble last month? Or the total production of ice cream in southern california in 2021? Or the frequency of tooth brushing in rural Idaho?
Block what you can is also about choosing the variables you have reason to believe have reasonably strong effects, you have a reasonable path to control, and there’s a reasonably ethical way to control them, randomizing the rest includes randomizing all the tooth brushing in Idaho and the motion of distant stars as well as things that you might even have theories of how they effect outcomes but which you have marginal control over or it would be unethical to alter… say ad-hoc usage of lotion among skin condition sufferers or OTC allergy sprays among hayfever patients or parental teaching around the dinner table among K-12 students, or the temperature and humidity on the day that paint is applied to buildings…
JYD: also yes, “the universe” can affect your random number generator. For example in Julia they use a generator called Xoshiro (derived from xor-shift-rotation). It’s entirely possible that while you’re generating your RNG numbers some cosmic rays flip bits in the generator state, and you get results that differ from what you’d get on a second run with the same seed. Of course, you can test that, even re-run multiple times to see that you get the same results. But, everything affects everything else, so why did you choose the particular seed? I often like to do something like using the date 20231113 or the first few digits of pi 3141591653 for example, but I could have chosen to hash the string representation of the date
julia> hash(“20231113-10:38:50”)
0xa98f3f8019c2b506
Or I could have chosen to take a picture of a tree in my backyard and run the resulting JPEG through sha256 and xored bits together to create a 64 bit seed… I could have done lots of things.
Fortunately, with a well designed PRNG every possible sequence is in some sense “equally likely” and mathematically speaking almost every possible sequence spreads the treatment and control groups evenly across any unknown variable you could imagine… this is just a property of high complexity sequences… essentially all of them are equally good. You can get “unlucky” but that’s extremely rare among choice of seeds or whatever. Maybe 1 in billions of trillions of seeds would give strong correlations with people’s big toe joint health. All the rest will spread that evenly among groups.
Typo in my pi digits! 3141592653 the universe is interceding to destroy my experiment! 😂
> “Block what you can” is the prescription to use all the known variables whether you have a theory for how they may influence the outcome or not.
You will have more blocks than elements to put into them as soon as you have more than a handful of variables – even sooner they are not binary.
Anyway, I may have misunderstood your point. Weren’t you questioning the value of the ‘randomizing what you cannot’ advice? In that case, what alternative advice are you putting forward?
Carlos’ point is really important to understand, if there are 20 variables you might care about and 5 different levels you’d divide them into, then 5^20 = 95367431640625 so the minimum sample size for any experiment is around 950 trillion samples since you’ll want some repetition within each group to get a sense of the variation.
Carlos,
Your earlier comment with the example of morning vs afternoon illustrates my point about too much being made about the value of randomization, leading folks to randomize on something they should be blocking on.
Yes, you are going to eventually exhaust your budget for experimental units but why does this mean that at some point using a RNG is better than just making a choice?
What is my recommendation?
#1 Don’t randomize on variables that can be blocked (or that otherwise could be used to allocate treatments in a way that would not admit potentially bad experimental designs).
#2 Randomize what you can’t, if you really feel the need to, but know that it doesn’t make a difference.
Daniel,
“ if there are 20 variables you might care about and 5 different levels you’d divide them into, then 5^20 = 95367431640625 so the minimum sample size for any experiment is around 950 trillion samples since you’ll want some repetition within each group to get a sense of the variation.”
This is a naive statement.
You really can only think of using a full factorial design here?
Why do you think you need ‘repetition’ to enable estimates of variation? You don’t.
Do you really think there might be a quartic polynomial that can be learned about among your ‘nuisance variables’? That seems a little extreme.
How does the number of such variables change the underlying question of whether using a RNG is better than just making a choice after everything that can be blocked has been blocked?
> why does this mean that at some point using a RNG is better than just making a choice?
What does it mean to “just make a choice”?
Is it fine to give the active treatment to patients who live with someone, or who live close to the hospital?
> #2 Randomize what you can’t, if you really feel the need to, but know that it doesn’t make a difference.
How do you know that?
JYD, randomization *is* “just making a choice” except it’s doing it in a way that has been validated through a massive test suite to correspond to the ideals of a high complexity sequence for most purposes.
There are too many ways to “just make a choice” which are influenced by human practical issues, like preferring to give treatments to the first people who come into the clinic, or to the children who seem to be in the most danger of severe illness, or to the wealthier people who bully you into it. Is it ok to “just ask the patient if they want the COVID vax or the placebo?” why is that not “just making a choice”?
What randomization is, is it’s a choice which is *defensibly justifying your assumption that it’s not confounded with anything important*.
If you’re doing Bayes, and you assume all of group A are exchangeable, the randomization process makes that assumption justified as opposed to simply naively optimistic.
Thank you JYD.
I study ADHD drugs by conducting mice experiments and my coworker’s keep insisting that I need to randomize the treatment assignment. But that requires me to first catch all the mice and mark them, then randomize, and then catch those that I have randomized to administer the drugs. I don’t see any value in it. Sure I could do it but I could also just pick some mice for the treatment group (as long as I don’t pick them based on any variable that I should block on).
From now on, I will just pick up 50% of the mice from my cage and use whatever mice I pick up as my treatment group. No need to catch them all initially. This saves me around three hours of work a week.
Rik:
There are designs other than complete randomization. You don’t need to catch all the mice at once, label them, and randomize them. Instead you could catch each mouse when you need it and flip a coin to decide which treatment to give it.
I may be misunderstanding, but wouldn’t that mean you’d be more likely to pick the mice that are easier to catch (slower moving, in worse shape) for the treatment? That seems bad.
Andrew, Somebody
I apologize, sarcasm dos not travel well over internet cables.
I do not study ADHD drugs.
And you are right Somebody.
My comment was just supposed to illustrate how dangerous it is to think that randomization “doesn’t make a difference” when you have blocked on all variables that you can, like JYD argues.
Rik:
Yes, I suspected you were joking—for one thing, researchers keep their mice in cages and don’t have to run around trying to catch them!—but I gave a straight response because it is an interesting point of experimental design, that it’s possible to randomize each item as it comes in, rather than needing to do all the treatment assignments at once.
It isn’t randomization that will produce an “unbiased” or “balanced” design it is the blocking. A particular design will inevitably be unbalanced on some variables whether or not we randomize.
Some might argue that randomization provides a process that in some infinite sequence of designs will in aggregate be balanced on all variables.
But given the general themes discussed on this site (e.g. nonstationarity) I doubt that argument is being made here.
And as a practical matter most experiments are themselves a ‘one off’ and so an appeal to some long run properties of a design process is of far less concern than doing the best we can with the current design.
Rik,
It seems like in your artificial example, you would want to block on a behavioral indicator you are suggestion might exist in your mice, rather than leave it to chance your randomization gives you a bad design.
Thanks for the links. I browsed to the the CQPI page, clicked the link to technical reports, and it took me to MINDS@UW which lists a bunch of technical reports, some by Box.
https://minds.wisconsin.edu/handle/1793/66926
I separately searched for Box on the MINDS@UW and found 35 reports…looks to me that someone was trying to move resourses like tech reports to a central site where they could be maintained.
I agree with MadisonMD that this is probably more mundane that any kind of active feud. I was at Madison in the early 2000s (as a grad student in the math department, not the stats department). At that time there was some other cross-department center, “Mathematical Systems Center” or something like that (I *definitely* have that name wrong). This center, which may have been some kind of successor to the Box-Hunter center, was also on its way out for the what I assumed were usual reasons: Initial funding rounds winding up, founding visionaries wandering off to retirement or other interests, and young faculty and grad students not seeing career paths through the center.
As a grad student, it’s not like I had any inside knowledge of any of this (so grains of salt all around). But it just felt like a normal entropy in the context of academic free-agency (and in light of a sort of seed-funding magical thinking that the NSF seemed to be into). That said, I agree Hunter’s lament: It is sad and wasteful to see so much effort going into building these things up only to have them atrophy and die after five or so years (and it’s galling for the cycle repeat itself).
There is a glorious promise in Box’s articles, such as those in “Box on Quality and Discovery” and “Statistics for Experimenters” which unfortunately is rarely fulfilled. He claims that, without dedicating your life to the study of Statistics, you can design experiments and analyse data well enough to improve areas to which Statistics has not previously been applied, and show that they have been improved. It is unfortunate that this dream has suffered both from hucksters fabricating statistical support for themselves and by the steady accumulation of ever more sophisticated techniques, which may provide some real gains, but which demand a much greater investment of study time and a much larger technical infrastructure in return.
The center was part of the Math Research Center located in the Warf building. MRC moved to Stanford in 1999 https://mrc.stanford.edu/about/about-us and the focus moved from Statistics to pure and applied Mathematics.
The helicopter experiment was an idea of Kip Rogers from Digital Equipment that Box popularized. Bill Hunter expanded on it https://williamghunter.net/articles/101-ways-to-design-an-experiment. We expanded on it some more https://www.tandfonline.com/doi/abs/10.1080/02664768700000027
The CQI reports are available in https://minds.wisconsin.edu/browse?rpp=20&offset=0&etal=-1&sort_by=-1&type=author&value=Box%2C+George&order=ASC
Seems however that this is not a complete list…
In 2023 I have an issue with the BH2 approach because in most cases you start an interaction with a “customer” by studying his data and not discussing an experimental set up, from scratch. The emphasis is on design for augmenting existing information which requires different methods….
Apropos of quality assurance I can’t resist mentioning my “discovery” in the course of a recent tour of the Guiness Brewery of Stella Cunliffe, who was chief statistician there and later president of the RSS. Her whole wikipedia article is inspiring, but I can’t resist quoting the following passage:
In this role, she developed important principles of experimental methods that are taught to this day. In the most famous example, she redesigned the instructions for quality control workers who were tasked to either accept or reject handmade beer barrels.[6] Before Cunliffe’s redesign, workers accepted barrels by rolling them downhill and rejected barrels by pushing them uphill, the more difficult task; thus, workers were biased to accept barrels even if they were flawed. Cunliffe redesigned the quality control work station so that it was equally easy to reject or accept a barrel, eliminating the prior bias and saving Guinness money in the process.
Daniel,
“There are too many ways to “just make a choice” which are influenced by human practical issues, like preferring to give treatments to the first people who come into the clinic, or to the children who seem to be in the most danger of severe illness, or to the wealthier people who bully you into it.”
Why the straw man?
These could all be blocking factors that don’t need to be randomized on.
What did you mean when you said “just having to make a choice after blocking as much as they can”?
Because when anyone suggest anything that could be a choice you say that it should have been a blocking factor instead.
Is there an opportunity for making arbitrary choices or not?
How do you assign the units in one block to alternative treatments?
Carlos,
“Because when anyone suggest anything that could be a choice you say that it should have been a blocking factor instead.”
Because three folks have now presented examples of variables that are measured on experimental units and that they are implying would have an association with the outcome variable of interest as a defense of randomization.
But this just illustrates my original point that making too much of the value of randomization, encourages people to stop blocking on things they should and randomize on them instead.
“How do you assign the units in one block to alternative treatments?”
In the strong sense, If you aren’t using information about your experimental units to make those assignments, then it shouldn’t matter whether you use an RNG or not.
Can you give an example of making a choice that’s not based on any information which is not random?