Postdoc and doctoral student positions in Bayesian workflow at Aalto, Finland

This job ad is by Aki

I’m looking for postdocs and doctoral students to work on Bayesian workflow. The candidates need to have knowledge of Bayesian inference and some experience with building models (for real applications, as part of methods development, or as part of courses). Although we have published Bayesian workflow book, there is still a lot more to do. The focus in the group is in cross-validation, model checking and inference diagnostics (see my publication list).

All positions are fully funded and the salaries at Aalto CS are 53k€-55k€ / year for postdocs and 40k€-45k€ / year for doctoral students. There are occupational healthcare and other benefits. Postdoc positions are typically offered for up to three years and doctoral student positions for four years. Starting dates are flexible and the details of each position will be agreed individually.

You can apply via joint ELLIS Institute Finland call and pick me as your preferred supervisor.

Bayesian Workflow free pdf!

Our wonderful new Bayesian Workflow book is now available as a free pdf! Just go the link—it’s right there!

I recommend getting the hard copy too because you’ll want to be able to read it while working on the computer, and the cost of the book is trivial compared to the benefit from faster learning that you will get by being able look at the book without taking up valuable screen real estate; also you can see connections when flipping through the pages that might not be apparent by viewing one page at a time on a screen.

Conversely, if you have the hard copy, you should still download the pdf because it fixes a bunch of minor errors that we caught after the book went to press. Also in the printed version we accidentally repeated some of the exercises in chapters 2 and 3. For the pdf we fixed this.

Regarding the content, as I wrote last month, with Bayesian Data Analysis, the big steps forward were:

  • Going beyond Bayesian inference to also consider Bayesian model building (as a researcher, you construct the model, it isn’t just given to you as in a textbook), model checking (breaking through the absolutely horrible attitude, common to Bayesians in the early 1990s, that the model was “subjective” and thus should not be checked), and model improvement (continuous model expansion, not the misguided idea of assigning posterior probabilities).
  • Going beyond simple conjugate models. BDA had lots of hierarchical models, also lots of computational tools so that you could fit the models you want by putting them together from understandable components. And I like how we had a clear separation between modeling and computing. The model comes first, then you figure out how to compute it. Or you set up a model that works within your computational constraints.
  • A Bayesian approach to sampling and causal inference. This was Rubin’s framework in which unobserved units in the population and unobserved causal outcomes are treated as missing data and are part of a joint probability model. We worked this out in chapter 7 of BDA (which became chapter 8 in the third edition of the book).
  • Lots of live examples. Not just “real-data examples,” but problems we’d directly worked on. This motivated us and I think it gave our readers a sense of how Bayesian methods worked not just in theory but in applied problems.
  • A pragmatic view of probability as a measurable quantity. That’s right there in chapter 1. Bayesian methods are not the product of a philosophical stance; they’re a way to connect models and data using probability.

I could go on and on, but for that I can refer you to the Bayesian Data Analysis book.

And these are the key innovations of Bayesian Workflow:

  • Going beyond Bayesian data analysis (model building, inference, model checking, and model expansion) to consider the larger process of statistical modeling, including comparisons of multiple models fit to a single dataset.
  • A fuller use of informative priors. This is a big deal. In BDA we still had a bit of the Bayesian cringe going on. One reason we’ve moved toward stronger priors is that the replication crisis has taught us that the amount of prior information available in any given problem is often approximately the same as the information coming from an experiment (see here, for example). Informative priors also fit our increased focus on generative modeling, and we’re doing a lot more prior predictive checking to understand the implications of our models.
  • More integration between modeling, data analysis, and computing. One way to see this is that the Bayesian Workflow webpage has the code to run all our examples. We also have lots of code snippets in the text as a way of demonstrating the way in which coding is central to our statistical workflow.
  • Lots more live examples. It’s been 30 years since BDA first came out. One reason that Bayesian Workflow has 11 authors is that different collaborators worked on different examples (but the three principal authors read through the entire book, so the general approach should remain coherent).
  • Simulation-based experimentation. This is something my colleagues have been doing more and more over the years. At its most basic, simulation-based experimentation provides a best-case baseline for statistical methods: if you can’t recover your quantities of interest with sufficient accuracy under ideal conditions (when your data are simulated from the model you’re fitting), then you know you’re in trouble. And often this is the case! Beyond that, we can simulate from one model and fit another, and see what happens. Simulation experiments aren’t always so easy to construct, as they involve specifying the entire data-generation process. But we think this is effort worth expending, as it involves thinking about the problem you’re working on.

I could go on and on, but for that I can refer you to the Bayesian Workflow book.

Here are the talks from StanCon 2026!

StanCon 2026 just happened!

And here are the talks:

Matthew Kay
Adrian Seyboldt
Paul-Christian Bürkner
Charles Margossian
Javier Enrique Aguilar
Sean Pinkney
Nikolas Siccha
Anna Dreber
Kaitlyn Johnson
Jonas Wallin
Pranav Sanke
Anna Elisabeth Riha
Colling Cademartori
Tim M. Szweczyk
Bob Carpenter
Chandler Ross
Nils Rudi
Fredrik Ronquist
Jakob Torgander
Aleksi Lahtinen
Ville Laitenen
Zeno Romero
Soham Mukherjee
Steve Bronder and Brian Ward
John Ashley Burgoyne

Lots of great stuff here. Check out the titles of the talks—all sorts of different topics.

Time to start planning for StanCon 2027.

The improvement in political analysis in the past 25 years, as demonstrated by excellent demonstrations of statistical workflow from Elliott Morris, Nate Silver, and Eli Mckown-Dawson

As with baseball, football, and basketball (and I’m sure other sports too), the standard of political analytics is just so much higher than it was, decades ago.

I was talking with Gustavo just the other day about Red State Blue State, and how that work was motivated by confusion following the 2000 election emanating from pundits of the left, right, and center. Back then I felt the compulsion to write a whole damn book to explain what was really going on. I even came up with an entirely new (to the best of my knowledge) concept, “second-order availability bias,” to explain how the journalists could’ve gotten things so wrong.

The concept of “second-order availability bias” never caught on, to say the least: it appears only once in the easily-accessed published literature:

So maybe it’s not such a useful psychological concept. What’s relevant here, though, is that the pundits were getting it so wrong, and with such a consensus, that I felt the need to refute them.

Nowadays, things are different. There aren’t so many all-purpose pundits like David Brooks—people who know essentially nothing and have no real interest in learning but present themselves as infallible experts—, and those who remain don’t have such a platform. Also, with partisan polarization, commentary has become fragmented, and we rarely see much of a consensus among pundits across the political polarization. It’s just not on the table.

Meanwhile, political analytics has become more and more impressive, with important contributions being made by academics, journalists, and political professionals. There’s still disagreement (as here) and some difficulties of communication (as here), but in the past two decades the level has gone up so much: the best analytics has become much more impressive, and what might be called replacement-level analytics has become much better too. I don’t think my own sophistication has increased much at all, and that’s one reason why I’m now less likely to crunch the numbers myself (as I did in the early morning hours of 5 Nov 2008) and more likely to just link to the analyses of others (as with Yair’s report on 2024).

Just today I came across two excellent examples online from journalist colleagues of mine.

Elliott Morris, “I re-analyzed the raw data from Wisconsin’s primary polls. Here’s what actually went wrong,” which features this split-the-difference summary that warms my Bayesian heart:
• Most of the miss in polls in Wisconsin is attributable to faulty demographic targets (too many young people). This inflated Hong’s vote margin by somewhere between 5 and 10 points.
• My best guess is that the race moved 6-10 points toward Crowley after pollsters released their final surveys.
• Non-ignorable non-response within demographic categories likely further inflated Hong’s vote margin by 2-5 points.
Morris goes through lots of details too. I haven’t tried to check any of this, but it seems reasonable. We’ve been saying for a long time that primary elections are hard to predict, but some polls are off by much worse than others, and it’s instructive to look into exactly how this can happen.

Beyond the details and the direct interest of this post to political organizations and pollsters, I appreciate Elliott’s work here because he goes beyond statistical generalities (“regression to the mean,” “sometimes you get a draw from the tail of the distribution,” etc.). This is an important statistical point: the “error term” is only an error term until you drill down, look at more data, and figure out what is going on. It’s so common for researchers to just take their numbers and not think about where they came from (as here)—and, indeed, academics and pundits alike can be rewarded for that sort of asinine don’t-look-carefully-at-the-data attitude. So it’s good to see Elliott demonstrating how it’s possible to do better—if you’re willing to put in the work.

Nate Silver and Eli Mckown-Dawson, “FLIPR 2026 midterm election forecast,” which leads Nate to summarize that “[Michigan Senate candidate] El-Sayed would be an underdog in an election held today and is an underdog in our “Lite” (polls-only) version. The fancy versions look at the fundamentals and are more convinced he’ll come back.”

What I really like about this is how “workflow” it feels. What I’m talking about here is how they fit two different models that are doing two different things, they learn something from the comparison, and then they track this back to their data. This sort of thing isn’t in the textbooks (well, it wasn’t until now) but it’s so important to good applied statistical work. So I love to see it here.

The point about these two posts, one by Morris and one by Silver and Mckown-Dawson, is not that they’re so amazing. I mean, yeah, they’re great, but the real point is how professional they are.

I remember Bill James once wrote, in reaction to the unexpected playoff heroics of Bucky Dent or Ray Knight or somebody like that, that, sure, it’s cool when someone steps up and does the unexpected, but what’s more impressive are those Eddie Murray types who can consistently deliver the expected. Because then you can design a game plan around them and not just have to hope for a miracle.

Morris, Silver, Mckown-Dawson, and others doing what’s now expected, doing it well, and demonstrating modern principles of statistical workflow . . . That’s impressive.

Frequentism for Bayesians: He wants to teach frequentist methods to engineering students with a strong Bayesian background

Beyond the teaching question, this is an interesting topic on its own: thinking about classical statistical ideas of point estimation, hypothesis testing, and uncertainty quantification, but taking Bayesian methods as a starting point.

The idea is that, instead of structuring this based on the methods of maximum likelihood estimation, null hypothesis significance testing, and confidence intervals, you start with the goals of estimating parameters, making predictions, comparing and evaluating models (i.e., alternative explanations of the world), and summarizing and working with uncertainty.

If you’re interested in using Bayesian methods, we have two books (Bayesian Data Analysis and Bayesian Workflow) on the topic. But with these methods under our belt, we can now go back to the motivating questions–the more general statistical goals, which exist without reference to any particular models or any particular Bayesian or frequentist methods–and consider them from scratch.

This seems important.

My discussion here is motivated by this question sent in by Opher Donchin:

Here’s one for your blog (although I’m happy to have your take as well).

I’m transitioning our undergrad biomedical engineering course this year from a standard frequentist syllabus to a Bayesian approach. We are mostly following the first 6 chapter of Bayesian Analysis with Python by Osvaldo Martin. In order to get the department to agree to this, I had to promise to teach basic frequentist methods.

Thus, I need to teach frequentist methods to students with a strong Bayesian background. There doesn’t seem to be much available on this. There is of course material comparing the two approaches, but I mean specific material designed to explain the frequentist approach to someone who knows the Bayesian approach.

I’d love to know if anyone knows of resources or has experience of insight or advice.

I replied that I’ll see what the blog commenters suggest, but in the meantime, I recommend chapter 4 of Bayesian Data Analysis as a start.

Donchin responded:

Yes. Chapter 4 is very good on the principles involved. Much of it addresses the Bayesian alternatives to frequentist procedures or the Bayesian perspective on them.

I’m wondering about something more concrete, aimed at a less sophisticated audience. That is, my students will know how to build models, how to interpret the posterior samples, the basics of Bayesian workflow and also model comparison. On the other hand, they will have no knowledge of confidence intervals, maximum likelihood estimates, hypothesis testing, or multiple comparison procedures. Reasonably, my department demands that they be able to read the biomedical literature where such terms are widespread.

I want to give them an understanding of frequentist procedures without getting bogged down in frequentist justifications.

To take an example, I want to explain what an F test calculates when understood within a Bayesian framework. To that end, I can show students a Bayesian model of a normal distribution of group means and a normal likelihood within each group. Then, I can work through what a Bayesian would need to calculate on that model to produce an F statistic and an F test. It has something to do with summaries of posterior distributions of ratios of variances.

What I’m hoping for is somewhere where such questions are worked through in detail so that it could be used for developing lectures.

Of course, chapter 4 of BDA is 10 years old at this point. I imagine some of the ideas may have developed since then.

Ahhh, good point! Chapter 4 of BDA is for statisticians who already know the classical methods and want to understand how these can be understood in light of Bayesian principles and adapted within a Bayesian workflow. But it’s really a completely different task to explain classical methods to students who haven’t already learned them. The idea would be to retcon ideas of classical statistics from a Bayesian angle.

This would be worth doing.

In the meantime, I recommend . . . chapter 4 of Regression and Other Stories, where we go through basic principles of point estimation, hypothesis testing, and uncertainty quantification from an applied perspectives. Also, if you flip through that book, you’ll see other places where we discuss classical procedures from first principles. We don’t have any F tests or multiple comparisons adjustments because I can’t bring myself to care about those things, but a lot else is there, so you might be able to put together much of what you need from that book.

In response to that recommendation, Donchin wrote:

I agree that Chapter 4 of Regression and Other Stories touches on many of the important points, but it is not sufficient for my needs.
I am trying to teach a Bayesian-first undergraduate statistics course to biomedical engineers that also gives a background in frequentist approaches allowing  them to function effectively in environments that require them to use or understand frequentist stats.
This means that the frequentist-realted topics we cover include:
  • Estimation: MLE, standard errors, confidence intervals
  • Hypothesis testing: p-values, Type I/II errors, power
  • Proportion tests and t-tests (independent and paired)
  • Effect sizes; multiple comparisons (briefly)
  • Linear and multiple regression: least squares, coefficient tests, R2, F-tests
  • ANOVA: categorical predictors, interactions, sums of squares, effect sizes
  • Pearson correlation and inference
  • Repeated-measures and mixed-effects models
  • Model comparison: AIC and cross-validation
  • Replication crisis / open science / pre-registration (not exactly frequentist, but still)
You can see the syllabus and the lectures in the student-facing version of the course repo at:https://github.com/opherdonchin/StatisticsCourse_36714361
Any further thoughts you might have would be great to hear.
Also, I would be happy to get this some visibility, in hopes that other people would be interested in providing feedback, using some of the material, or just doing it better.

I don’t think I could bring myself to teach a lot of the above topics, except in an “inoculation” sort of way, but I recognize that many students will need it, so if anyone has some good suggestions for Donchin, just leave them here in the comments!

7 1/2 Schools—mistakes were made

Mitzi Morris sends along the paper trail (I love archaic idioms) for her data spelunking in advance of her all-day introductory Stan tutorial next week at StanCon (we’ll see you in Uppsala). She was trying to reproduce the results from my decade-old case study on hierarchical modeling when she ran into a couple discrepancies.

The case study was designed to provide a Bayesian replication of the results in Efron and Morris’s five decade old paper on hierarchical modeling (aka Stein’s estimator, aka population regularization, aka “empirical” Bayes), which is still under a paywall courtesy of our “friends” at the American Statistical Association. If your organization isn’t paying ASA for access to a paper that an academic donated for free 50 years ago, I’ll leave you to find your own pirated copy in good conscience, or you can follow the link and let Google hoist the Jolly Roger for you (now that’s an even more obscure and archaic reference).

My case study’s been out for ten years. Mitzi found a problem when matching the data provided in the R package pscl against that in Efron and Morris’s paper. In particular, the data for the player named “Williams” was wrong. Efron and Morris manually transcribed the data from a newspaper with the goal of finding a bunch of players with the exact same number of at bats on a given day (45, it turns out). They did so accurately.

Sadly, I imported an error in the 18-item data set from the R package. Here’s an edited report from Claude after it tracked down what went wrong.

  1. The 1975 paper is correct, although it only gives the surname “Williams.”

  2. The R package pscl expanded that name incorrectly to Billy Williams, while retaining Walt’s team and remaining-season data. Published pscl data.

  3. The Bean Machine tutorial () contains the exact mixed row found in your CSV:
      Billy,Williams,...,591,0.3299,636,205,0.251
    

  4. That row combines Billy Williams’s at-bats and hits with Walt Williams’s .251 average. Thus, ID 15 in your data should be corrected throughout to:
      Walt,Williams,45,10,0.222,270,0.2556,315,79,0.251
    

But wait, that’s not all. There was another typo in the Bean Machine tutorial. They mistranscibed Ron Swoboda‘s surname as “Swaboda.” I should have caught that as I was a huge baseball fan in the early 1970s, am an excellent speller, and grew up around a lot of Poles and Czechs in Detroit and its western suburbs. Mitzi then had the following conversation with Claude after she noticed the typo. This may be the first time I’ve seen Claude crack a spontaneous joke.

Mitzi: did that [the typo] come also from the R file?

Claude: If you mean the Swaboda misspelling: no. The pscl R dataset correctly uses Ron Swoboda. The misspelling appears in the Bean Machine tutorial
(https://beanmachine.org/docs/overview/tutorials/Hierarchical_modeling/HierarchicalModeling/), which contains Ron,Swaboda.

Mitzi: are you sure? I’ve heard it both ways.

Claude: Yes—despite what Shawn Spencer might say, it is definitively Ron Swoboda, with an o.

Mitzi: who’s Shawn Spencer? in 2017, Bob Carpenter published this dataset in a Stan case study.

Claude: Shawn Spencer is the fictional detective in the TV comedy Psych, played by James Roday Rodriguez. “I’ve heard it both ways” is his recurring response when someone corrects him—hence my reference.

Mitzi and I use “I’ve heard it both ways” all the time, having watched umpteen seasons of Psych. I would recommend if you want a lighthearted detective story that doesn’t take itself too seriously with lovable characters and you can put up with the ubiquitous mid-00s background sexism and racism.

Bayesian computation meeting in France in May, 2027

OK, this one’s time constrained so I’ll post it right away, not on the usual lag.

Christian Robert points to this conference announcement:

We invite proposals for the BayesComp 2027 mirror event in Aussois, France, that will take place on 17-21 May, 2027. The event will complement the main BayesComp 2027 meeting in Texas and provide a European forum for researchers in Bayesian computation and related areas. Plenary sessions will be broadcasted from Texas, while talks in Aussois will be recorded.

Please submit one form per proposed contribution. Proposals may be for contributed talks, or poster presentations. The programme committee will review submissions based on scientific quality, timeliness, coherence, and expected appeal to the BayesComp community. Unless the submission mentions otherwise, talk proposals that are rejected will be automatically added as posters.

Details about the venue (CAES Paul Langevin), potential registration fees, and technical details are being finalized and will be announced on the event website.

Looks like fun! And Bayesian computation is important.

Survey Statistics: wanting workflow

Last week Andrew commented that we need a more transparent workflow for survey statistics. So I looked in the new Bayesian Workflow book:

Chapter 19 “Building up to a hierarchical model: Coronavirus testing” is a case study about a 2020 survey that tested n = 3330 residents of Santa Clara County, California for SARS-CoV-2 antibodies (Bendavid et al. 2020a, 2020b). y = 50 people tested positive.

To estimate population prevalence, they want to account for measurement error in the test (“Measurement“) and differences between sample and population (“Representation“), both sides of Groves et al. Figure 2.5:

Image

The measurement error model includes specificity gamma = P[test negative | no disease] and sensitivity delta = P[test positive | disease], which take population prevalence pi to test positivity rate p:

Bendavid et al. (2020a) analyzed their data using gamma = 0.995 and delta = 0.80. But these aren’t known exactly, so Gelman and Carpenter 2020 recommended using priors from previous studies to reflect uncertainty:

They extend these priors to account for multiple studies. They look at sensitivity to these priors (overloaded term here: “sensitivity”). Cool stuff !

In this Survey Statistics series we hadn’t yet talked about models for measurement error in the outcome variable y. We’ve seen measurement error in an adjustment variable X, e.g. recalled vote (“is a mismeasured X better than none at all ?”, “more adventures in mismeasured X”, “more on recalled vote”, “it is (still) the people”).

On the “Representation” side, Bendavid et al. (2020a, 2020b) adjusted for differences between sample and population using weights and encountered the difficulty we saw last week: adjusting for lots of variables can lead to very large weights. Maybe modeling can help (see last week’s “structured MRP to smooth survey weights”).

So Gelman and Carpenter 2020 proposed replacing (19.2) above with

y_i ~ bernoulli(p_i)
p_i = (1-gamma) (1 - pi_i) + delta pi_i

and using Multilevel Regression and Poststratification (MRP):

I thought this story sounded familiar and remembered Andrew talked about it at his Birthday workshop‘s closing talk (this part on YouTube, minute 2 to minute 8).

Andrew said that when his friend tried the method from Gelman and Carpenter 2020 it “crashed and burned” ! There is a new paper about this: Kuh, Kennedy, Chen, and Gelman 2026. They present “a statistical workflow for diagnosing unexpected results in complex models”. These authors have researched a lot about evaluating MRP models (see “should MRP workflow include LOCO-CV ?”). We will have to revisit their new paper later in this series.

The blessing of dimensionality

Jonathan “No Trump” Falk writes:

Your post this week on Integration and Differentiation got me thinking, never a good sign. Adding to this, I only really learned in January what Attention means, the foundation of LLM. And what it means is that in a vector space of 14,000 dimensions or so, you can express all manner of nuance… enough to start vitriolic arguments about the humanity of the output.

As I started thinking about this, and as your post crystallized, I have spent 50 years fighting the Curse of Dimensionality. I know this curse in my marrow. Brilliant inferences await me, but the space in which these insights are found is simply too vast to explore. So we simplify, reducing the dimensionality to something that while still vast, is confined to a hyperplane where we can, like Plato, see the projections of truth, not the truth itself.

But then what LLMs and their generation have taught me is that nuance, which is really just the inverse of inference (in that it’s the vast set of all things consistent with some inference) has an amazing boon of dimensionality. There appears to be no thought that can’t be described by a 14,000 dimension or so vector whose tuning has the huge advantage that 14,000-dimensional space is so empty that tiny nuances can be readily distinguished in such a space, so that you can hide uniqueness in the vastness of 14000-dimensional space that you couldn’t recover in a raw search in that same space.

I’m sure this inversion of the Curse of dimensional search into the Boon of nuance in high dimensions is not original to me, but I think it’s worth noting, and your post was the impetus.

I replied by pointing to our of our very earliest blog posts, The blessing of dimensionality, where I wrote:

The phrase “curse of dimensionality” has many meanings (with 18800 references, it loses to “bayesian statistics” in a googlefight, but by less than a factor of 3). In numerical analysis it refers to the difficulty of performing high-dimensional numerical integrals.

But I am bothered when people apply the phrase “curse of dimensionality” to statistical inference.

In statistics, “curse of dimensionality” is often used to refer to the difficulty of fitting a model when many possible predictors are available. But this expression bothers me, because more predictors is more data, and it should not be a “curse” to have more data. Maybe in practice it’s a curse to have more data (just as, in practice, giving people too much good food can make them fat), but “curse” seems a little strong.

With multilevel modeling, there is no curse of dimensionality. When many measurements are taken on each observation, these measurements can themselves be grouped. Having more measurements in a group gives us more data to estimate group-level parameters (such as the standard deviation of the group effects and also coefficients for group-level predictors, if available).

In all the realistic “curse of dimensionality” problems I’ve seen, the dimensions–the predictors–have a structure. The data don’t sit in an abstract K-dimensional space; they are units with K measurements that have names, orderings, etc.

For example, Marina gave us an example in the seminar the other day where the predictors were the values of a spectrum at 100 different wavelengths. The 100 wavelengths are ordered. Certainly it is better to have 100 than 50, and it would be better to have 50 than 10. (This is not a criticism of Marina’s method, I’m just using it as a handy example.)

For an analogous problem: 20 years ago in Bayesian statistics, there was a lot of struggle to develop noninformative prior distributions for highly multivariate problems. Eventually this line of research dwindled because people realized that when many variables are floating around, they will be modeled hierarchically, so that the burden of noninformativity shifts to the far less numerous hyperparameters. And, in fact, when the number of variables in a a group is larger, these hyperparameters are easier to estimate.

I’m not saying the problem is trivial or even easy; there’s a lot of work to be done to spend this blessing wisely.

It’s been over 20 years so worth sharing the point again.

Survey Statistics: structured MRP to smooth survey weights

Last week, Raphael K shared a concern: adjusting for lots of variables can lead to very large weights. So today let’s dive into Si et al. 2020, who saw this in constructing survey weights for the NYC Longitudinal Study of Wellbeing.

To adjust for lots of variables, Si et al. 2020 turned to MRP (Multilevel Regression and Poststratification) and equivalent weights based on these models (see “struggles with equivalent weights”continued struggles, and “equivalent models, equivalent weights (locally)”).

In a simulation study they compare:

  • Ind-P: MRP with commonly-used Independent Normal priors
  • Str-P: MRP with a structured prior, see below.
  • Ind-W: equivalent weights version of Ind-P
  • Str-W: equivalent weights version of Str-P
  • Rake-W: classical raking weights, a type of calibrated weights
  • PS-W: classical poststratification weights, another type of calibrated weights
  • IP-W: inverse probability of selection weights

They cover the “3 flavors of survey weights”: equivalent weights, calibrated weights, IP-W.

I won’t bury the lead, they found MRP performed best, then equivalent weights, then classical weights (calibrated or IP-W). See their Figure 4.1 for the simulation scenario without terribly many empty poststratification cells:

With many empty poststratification cells, Str-P outperforms Ind-P. (They don’t redo Figure 4.1 for this scenario, which confused me a bit.) So what is this structure that helps ?

In “improving with structure” we saw that Gao et al. 2021 found it helpful to use the ordinal structure of variables like age. Si et al. 2020 use the interaction structure:

We induce structured prior distributions to be able to handle deep interactions and account for their hierarchy structure, where the high-order interaction terms will be excluded if one of the corresponding main effects is not selected.

I asked about sparse priors for MRP back in “Sparsified MRP”. I didn’t remember that Si et al. 2020 had worked on this ! Ok so they write their structure more generally but I find it easier to read with a specific example. Consider just 2 variables from their motivating NYC Longitudinal Study of Wellbeing: age (5 categories) and race (5 categories). Here’s how their Ind-P prior differs from the Str-P:

(I had Claude type up my hand-drawn notes, though I still share Brendan Leonard’s preference for hand-drawn materials.)

Si et al. 2020 say these are similar to the Horseshoe prior. It differs in two ways, I think ? First, Si et al. 2020 have the selection at the batch level (e.g. age or race). The usual Horseshoe would have local scale lambdas for each age and race category. Second, the usual Horseshoe would use a half-Cauchy rather than half-Normal prior on these local scales.

Ok let’s get back to the original concern: adjusting for lots of variables can lead to very large weights. Si et al. 2020 show in Figure 5.1 that the equivalent weights based on this structured prior model look much less variable than calibration weights (I don’t see the IP-W weights in the figure itself):

 

“Placebo tests deserve a model, not just a glance.”

Miha Gazvoda shares this post with the above title and the subtitle, “Using Bayesian multilevel models to correct bias and calibrate uncertainty.”

He’s using the chickens model from our Slamming the Sham paper in the more general setting of placebo control tests.

In econometrics, a “placebo control test” does not need to literally involve a placebo treatment; it more generally is used to describe a procedure in which the same statistical analysis that was used to estimate a causal effect is applied to a different dataset, or a different part of the existing dataset, in which the treatment did not occur.

For example, if you want to measure the effect of an intervention that occurred in 2021, you could repeat the analysis but using data from a different year. Or if you want to measure the effect on a particular outcome, you could repeat the analysis but looking at a different outcome that should be unaffected by the treatment.

The idea is that, if your estimation method has an artifact or systematic bias, this should show up in the placebo analysis as well, indicating a problem. Conversely, if the placebo analysis does not show an effect, this is taken as evidence that there is no artifact.

In his post, Gazvoda argues that, rather than using the result of the placebo check to make a go/no-go decision, it should be possible to partially adjust the treatment effect to account for the information in that supplementary analysis.

This makes sense to me, and of course I’m happy that he’s using our chickens model.

There’s a tricky thing going on here with the placebo check, which, interestingly, arose in the chicken example too, and that is that we usually don’t have any good theory for where the effect is coming from in the placebo control. After all, if our causal identification is working as designed, we shouldn’t even need the placebo comparison, as we’re already getting an unbiased estimate of the treatment effect. The placebo control is typically there to address unspecified concerns of bias. And, indeed, in practice, researchers don’t always do placebo controls. And when, as hoped, the placebo control shows no statistically significant effect, the inclination is to take that as a reassurance and move on, in the same way that is done with other robustness checks.

The lesson Gasvoda takes from the chickens example is that if replications are available, you can assess the evidence for the placebo adjustments being relevant: you can fit a multilevel model to estimate how much adjustment needs to be done.

From a sociology-of-science point of view, it’s interesting to me that the conventions in biomedical statistics and econometrics go in opposite directions:

– In biomedical statistics, the default recommended behavior is to compute the difference in differences, taking the estimated effect from the main experiment and subtracting the estimate from the placebo experiment. As we explain in the chickens paper, this correction has the disadvantage of doubling the variance of the estimate, a true statistical crime if, as is often the case, the effect of the placebo treatment is indistinguishable from zero.

– In econometrics, the default procedure, if you’re calling it “difference-in-differences estimation,” is the same as above. But if you’re calling it “placebo control,” and the placebo estimate is not statistically significant from zero, the default is to ignore the placebo results entirely, not to adjust for them.

In general we recommend a partial adjustment, with the amount of adjustment depending on the problem at hand. If there is internal replication, as in the chickens example, the appropriate adjustment factor can be estimated from the data. If it’s a one-shot experiment, you’ll need to use prior information. I don’t have any good examples demonstrating how to do that; it’s something we should do.

Posterior predictive checking is for non-Bayesians too!

When I first started working on posterior predictive checking back in 1988, it was as a device for determining equivalent degrees of freedom for a chi-squared test for a model with constrained parameters–in that case, positivity restrictions in an image reconstruction problem; see here for background.

The idea is that the distribution of the test statistic depends on the true parameter vector–there is no simple pivotal quantity as would exist under a linear model with no constraints–but you can work out the distribution conditional on the true parameter vector, and then you can average this distribution over the posterior for the parameters.

You can think of this as Bayesian–the marginal posterior distribution of the test statistic–but at the time I was thinking of it more as a generalization of the existing “plug-in” approach that would use a point estimate of the parameter vector. From that perspective, posterior averaging is a technique for getting a distribution with better frequency properties than you’d get from the maximum likelihood estimate, and that’s because in an image reconstruction problem with positivity constraints and lots and lots of pixels, the maximum likelihood estimate will almost certainly be on the boundary of parameter space. So just about any sort of averaging should get you closer to the true parameter value.

As noted above, I started working in this area in 1988. Around 1991 I got my thoughts organized enough to write a paper and give a talk on the topic, and . . . it got a generally hostile reception! The Bayesians didn’t like it because they didn’t like chi-squared tests or frequentist hypothesis testing more generally. They wanted me to do Bayes factors, which even then I realized were generally a bad idea (for more on the topic, see chapters 6 and 7 of BDA3 (it was all in chapter 6 of the first two editions) or this article from 1995). The non-Bayesians didn’t like it because it was Bayesian, and because I just presented the method, without trying to justify it based on minimax properties or whatever.

Xiao-Li, Hal, and I finally published a version of the article a few years later, and I resigned myself to the fact that posterior predictive checking was going to remain in the Bayesian world. Predictive checking was a hard sell to the Bayesians, but, after a few decades of exposure to BDA and lots of applied examples, they started to get used to the idea. I gave up on trying to promote the idea outside the Bayesian community.

But . . . it turns out that posterior predictive checking has been taken up for non-Bayesian uses! Aki pointed me to this R package called “performance” that “provides posterior predictive check methods for a variety of frequentist models.” Here’s their vignette on checking model assumption – linear models,” and here’s the article, performance: An R Package for Assessment, Comparison and Testing of Statistical Models, by Daniel Lüdecke, Mattan Ben-Shachar, Indrajeet Patil, Philip Waggoner, and Dominique Makowski. It came out in 2021 and has already been cited over 2500 times.

So it looks like people really are using posterior predictive checks for non-Bayesian models. And it took less than 40 years for it to happen!

“Over-coverage caught by pre-registration: 47 of 56 inside a stated 50% interval”

Alex Malinowski has a question about evaluating the calibration of interval forecasts:

We publish interval forecasts under a pre-registration scheme: each forecast is serialised, hashed and timestamped into a Bitcoin block before publication, so the stated interval cannot be adjusted after the outcome is known. Across 56 resolved forecasts our stated 50% intervals contained 47 outcomes; at the 24-hour horizon, 36 of 44. Under honest 50% intervals, P(>=36 of 44) is about 1.3e-05.

The diagnosis, and the part I would most like criticised: walk-forward across 20,058 observations showed the miscalibration was conditional rather than uniform – 55.6% coverage across all days, 67.1% restricted to the calmest fifth. Our first hypothesis, that the sample carried too much old high-volatility history, was tested and rejected. The surviving explanation is that interval width ignored the current volatility regime. Conditioning the historical sample on regime measured at each window’s open (not its close, which leaks the outcome) brings coverage to 50.8%.

Two things I am unsure about. The instruments are correlated, so effective sample size is well below 56 and I have not done that properly. And at a 30-day horizon the conditional interval comes out wider than the unconditional one, which I have kept but cannot fully account for.

I asked him what was the application, and he replied:

Crypto prices. 24-hour and 7-day intervals on eight pairs (BTC, ETH, SOL, BNB, XRP, DOGE, ADA, LINK), scored against Binance closes.

I chose it as the substrate rather than the subject. Outcomes resolve within a day, the reference price is unambiguous, and there is no data-vintage or revision problem, so a coverage check accumulates evidence quickly and cheaply. The same construction is what we use for sports and prediction-market questions, but those resolve slowly and the sample is thin.

I realise crypto invites a certain reaction, and the reaction is mostly deserved. The calibration question doesn’t depend on it – the same test applies to any published range.

He adds:

If it helps frame it, the two things I’m least confident about are the ones I’d want readers to attack:

1. The eight instruments are correlated, so the effective sample size is well below 56 and I have not handled that properly. The binomial p-value I quoted assumes independence it doesn’t have.

2. At a 30-day horizon the regime-conditional interval comes out wider than the unconditional one. I kept the result because it’s inconvenient, but I can’t fully account for it beyond “calm periods have historically preceded larger monthly moves”, which feels like a description rather than an explanation.

Data and outcomes are CC-BY, one row per sealed forecast with its hash:
https://huggingface.co/datasets/neuportal/neuportal-sealed-crypto-forecasts

My only quick thought here is that it’s indeed a good idea to look at conditional calibration, but you have to be careful only to condition on things in the forecast, not on the outcome. So when he writes, “67.1% restricted to the calmest fifth,” this is relevant as long as “calmest fifth” is something that can be determined before the outcome occurs.

More generally, yeah, it’s hard to evaluate the calibration of forecasts for correlated outcomes–this comes up in election prediction too. My only answer is to try to model intermediate outcomes as much as possible to avoid getting stuck in a purely empirical mode of evaluations.

Maybe you in comments have other ideas?

Reviews of our Bayesian Workflow book from Bin Yu, David Spiegelhalter, Brad Efron, Christian Robert, Rohan Alexander, and Mine Doğucu!

Roughly speaking, Bayesian Workflow is to Bayesian Data Analysis in 2026 what Bayesian Data Analysis was to earlier Bayesian books in 1995: it builds upon everything that came before.

With Bayesian Data Analysis, the big steps forward were:

  • Going beyond Bayesian inference to also consider Bayesian model building (as a researcher, you construct the model, it isn’t just given to you as in a textbook), model checking (breaking through the absolutely horrible attitude, common to Bayesians in the early 1990s, that the model was “subjective” and thus should not be checked), and model improvement (continuous model expansion, not the misguided idea of assigning posterior probabilities).
  • Going beyond simple conjugate models. BDA had lots of hierarchical models, also lots of computational tools so that you could fit the models you want by putting them together from understandable components. And I like how we had a clear separation between modeling and computing. The model comes first, then you figure out how to compute it. Or you set up a model that works within your computational constraints.
  • A Bayesian approach to sampling and causal inference. This was Rubin’s framework in which unobserved units in the population and unobserved causal outcomes are treated as missing data and are part of a joint probability model. We worked this out in chapter 7 of BDA (which became chapter 8 in the third edition of the book).
  • Lots of live examples. Not just “real-data examples,” but problems we’d directly worked on. This motivated us and I think it gave our readers a sense of how Bayesian methods worked not just in theory but in applied problems.
  • A pragmatic view of probability as a measurable quantity. That’s right there in chapter 1. Bayesian methods are not the product of a philosophical stance; they’re a way to connect models and data using probability.

I could go on and on, but for that I can refer you to the Bayesian Data Analysis book.

And these are the key innovations of Bayesian Workflow:

  • Going beyond Bayesian data analysis (model building, inference, model checking, and model expansion) to consider the larger process of statistical modeling, including comparisons of multiple models fit to a single dataset.
  • A fuller use of informative priors. This is a big deal. In BDA we still had a bit of the Bayesian cringe going on. One reason we’ve moved toward stronger priors is that the replication crisis has taught us that the amount of prior information available in any given problem is often approximately the same as the information coming from an experiment (see here, for example). Informative priors also fit our increased focus on generative modeling, and we’re doing a lot more prior predictive checking to understand the implications of our models.
  • More integration between modeling, data analysis, and computing. One way to see this is that the Bayesian Workflow webpage has the code to run all our examples. We also have lots of code snippets in the text as a way of demonstrating the way in which coding is central to our statistical workflow.
  • Lots more live examples. It’s been 30 years since BDA first came out. One reason that Bayesian Workflow has 11 authors is that different collaborators worked on different examples (but the three principal authors read through the entire book, so the general approach should remain coherent).
  • Simulation-based experimentation. This is something my colleagues have been doing more and more over the years. At its most basic, simulation-based experimentation provides a best-case baseline for statistical methods: if you can’t recover your quantities of interest with sufficient accuracy under ideal conditions (when your data are simulated from the model you’re fitting), then you know you’re in trouble. And often this is the case! Beyond that, we can simulate from one model and fit another, and see what happens. Simulation experiments aren’t always so easy to construct, as they involve specifying the entire data-generation process. But we think this is effort worth expending, as it involves thinking about the problem you’re working on.

I could go on and on, but for that I can refer you to the Bayesian Workflow book.

And now for the reviews

But you don’t have to trust me on this! Just listen to some of the eminent statisticians and educators who’ve reviewed our book:

Bin Yu (University of California):

An outstanding, protocol-driven guide for Bayesian data analysis, Bayesian Workflow by Gelman, Vehtari, McElreath and co-authors delivers a practical and comprehensive framework for iterative modeling, emphasizing simulation, diagnostic checks, and rigorous empirical validation, and with a long and impressive list of case studies. By treating data analysis as a structured, verifiable workflow, it provides an indispensable toolkit for diagnosing model failures, refining priors, and building reliable data analysis systems for reproducible conclusions, useful for beginning and veteran data analysts alike.

David Spiegelhalter (Cambridge University):

This is not a typical methods textbook, but instead it guides the reader through the whole process of fitting, critiquing and adapting statistical models to real-world problems. It is full of the accumulated wisdom of skilled practitioners, teaching through demonstration rather than theory, with both basic and highly sophisticated examples. I strongly recommend this book to statisticians who really want to understand what they can learn from their data.

Brad Efron (Stanford University):

A bravura performance…Gelman, Vehtari, McElreath and friends develop in detail a practical Bayesian data analysis workflow, from acquisition to final report, including full computational guidance.

Christian Robert (Université Paris Dauphine):

This original, thought-provoking, and transformative book is much much more than an implementation manual for Bayesian Data Analysis, even though it shares almost the same perspective. (The first sentence of the book states that the authors’ “conceptions of statistical practice, and of Bayesian statistics, have changed over the years”.) By providing a modus vivendi for undertaking Bayesian modelling from scratch in realistic settings where models are not magicked out of the blue, the authors explicit and rationalise the many steps required by such a bottom-up modelling protocol (“not a checklist, not a cookbook”, and not a flowchart!) in real situations. The contents read very well and very smoothly, with a seamless conjunction of intuition, modelling advices, computational details, and comparison tools. While unsurprisingly Bayesian, the perspective adopted therein remains both open and inclusive, with a welcome humility about the limitations and challenges of Bayesian workflows. This book should thus appeal to and profit a wide variety of readers, as providing guidance through an extensive collection of highly detailed examples, with shared code and exercises.

Rohan Alexander (University of Toronto):

Some statistics books show you how to beat an egg, others are recipe books: if this, then that style. This book teaches you how to cook. Written by authors who established so much of how we do Bayesian statistics, this new book is an indispensable guide for analyzing data in a trustworthy way. It walks you through the actual steps involved in building models to explore and understand datasets. Part 4 is particularly excellent – the authors provide many end-to-end case studies that will be useful for both practitioners and students. It highlights the value of their workflow-based approach. Filled with chatty asides, the book introduces the Bayesian workflow to a broad audience. It embraces the frustrations and complexities of actually doing Bayesian statistics and provides specific guidance throughout. Each chapter contains exercises and it could be the basis of an upper-year undergraduate course, or a first-year grad course, in applied statistics. It will be used for many years to come.

Mine Doğucu (Harvard University):

What makes Bayesian Workflow so exceptional is how it seamlessly pairs profound ideas about modeling with the adoption of modern computational practice. By centering the messy, iterative process of modeling through real-world case studies, the authors reject rigid cookbooks and checklists in favor of building deep situational awareness. Because the ideas are so clearly articulated and deeply applied, this book serves as an invaluable pedagogical resource. With its practical exercises, individual chapters or the text as a whole can easily be integrated into upper-level undergraduate or graduate courses, while also remaining accessible for self-guided readers. It is an indispensable read for anyone with foundational knowledge in Bayesian methods, regardless of whether they are applied practitioners, software developers, or methodologists.

You might also be interested in the journal issue on statistical workflow that we recently edited for the Philosophical Transactions of the Royal Society.

Again, here’s Bayesian Workflow on Amazon, here’s the publisher’s website, and here’s our website with data, code, and lots more.

Enjoy.

Survey Statistics: quantifying uncertainty in ranked choice voting polls

We’ve talked about uncertainty in polls (see Margin of Error, Total Margin of Error, Total Margin of Error II) and we’ve talked about ranked data (see exploded logit !). A new paper, Rosenman & Liang 2026, looks at uncertainty in ranked choice voting (RCV) polls.

Recall the multinomial logit model that Train (2009) Chapter 7 calls the exploded logit:

P[ranking Other then Left then Right] = exp(f_Other) / sum_c’ exp(f_c’)   *   exp(f_Left) / (exp(f_Left) + exp(f_Right))

Without covariates, it has only 3 parameters: f_Other, f_Left, f_Right. It makes the independence from irrelevant alternatives (IIA) assumption to go from these 3 parameters to rank probabilities.

In contrast, the multinomial model in Rosenman & Liang 2026 does not make the IIA assumption and has 14 parameters, one for each of 15 possible rankings minus one so they sum to 1:

P[ranking Other then Left then Right] = pi_{Other, Left, Right}

Rosenman & Liang 2026 note that in RCV the election outcome is not expressable as one parameter. Instead, the winner is determined by instant runoff:

  1. If a candidate wins >50% of first choice votes, they win.
  2. Otherwise, the candidate with the least first choice votes is eliminated, and each ballot counts for its top remaining choice. Return to step 1.

Say you use polling data to estimate rank probabilities pi_j for each ranking j. These estimates differ from the true probabilities due to many sources of error (see our favorite Figure 2.5 from Groves et al. shown in quantity vs quality and is a mismeasured X better than none at all ?). Rosenman & Liang 2026 focus on sampling error.

How can we propagate uncertainty about the rank probabilities pi_j to uncertainty about the RCV winner ? If you have draws from the posterior of pi_j, you can do instant runoff on each to get a winner for that draw. This gives win probabilities according to your model and data.

To see the importance of uncertainty in RCV, let’s look at their 2022 Alaska House special election example. With 3 candidates, RCV is determined by 5 margins (see their Lemma 1). Most of these margins are well-identified by the data, but 2 were quite close: Palin vs Begich first choice margin and Peltola vs Palin pairwise margin. They plot these 2 margins in the right panel of Figure 1. The true outcome is the black dot, with sampling uncertainty shown as ellipses around it. For small sample sizes (the biggest ellipse), we see that a plurality of the mass falls into green, where point estimates would declare that Begich wins. Uncertainty quantification would help put this in context, giving all candidates win probabilities around 20-40%, showing the race is difficult to call with such small data.

For details, see Rosenman & Liang 2026.

 

“Making Statistics Work: Information Theory and Bayesian Inference”

I took a look at the above-titled book by economists Duncan Foley and Ellis Scharfenaker. It’s an interesting read, in many ways a throwback to the 1950s when a group of mathematicians brewed a heady mix of operations research, game theory, probability theory, and economics in an attempt to create a unified theory of social science, or to map the limitations of this effort. Important figures in this effort include John Maynard Keynes, John Von Neumann, Jimmie Savage, Milton Friedman, Duncan Luce, Howard Raiffa, Kenneth Arrow, Herbert Simon, Ed Jaynes, . . . a whole bunch of people who are still remembered today.

Back in the day, Bayes was seen alternately as Jesus or the Devil, and there were hopes of a grand synthesis of subjective probability and local information in markets, a connection between formal statistical inference and individual decision making.

In retrospect, the cognitive science of the 1950s wasn’t all there, and hierarchical modeling hadn’t been integrated into Bayesian inference. Also some key pieces such as posterior predictive checking and general-purpose Bayesian computing weren’t there. So any attempts at unification were premature. Not that it was a bad idea to try! Much is learned from incomplete efforts. It’s just clear in retrospect that any unified theories of the time were bound to fail.

From the perspective of seventy years later, we can see Bayesian inference as a useful part of the statistical toolkit, a way to place regularization (a central part of all modern machine learning) in the context of scientific modeling. Many problems that can be solved with an entirely Bayesian approach, and others can be viewed as approximate Bayes–or, to put it another way, Bayesian ideas can help with all sorts of statistical modeling problems, even when other inferential methods are used.

What I’m saying is, it’s a good idea for everyone doing statistics or machine learning to understand the basics of Bayesian inference and computation, prior and predictive checking, and Bayesian model expansion, for their own sake and also as a way to make sense of statistical learning.

“Making Statistics Work: Information Theory and Bayesian Inference” is one of the most unusual statistics books I’ve ever read. I don’t agree with much of it–for example, right on the second page they start talking about “prior beliefs,” which isn’t how I think of things at all (see here and here)–and it’s written in a mathematical style which seems old-fashioned to me but has a kind of charm. You could almost say that it’s the statistics book that William Feller would’ve written had he been converted to Bayesianism.

What I really like is that the book is what it is–an forthright attempt at a modern expression of the aimed 1950s synthesis of mathematical statistics, physics, and economics. It clocks in at a crisp 300 pages. It has zero overlap with Bayesian Data Analysis and Bayesian Workflow, and that’s just fine. As I said, I don’t really buy their synthesis myself, but I respect their attempt. You can judge it as you will.

The high cost of split R-hat

This post is by Bob.

I’ve been thinking a lot lately about R-hat given that I’m using it for online converging monitoring in our new Walnuts implementation. In that setting, where I use Welford accumulators to update R-hat estimates every iteration, I can’t use split R-hat without way too much buffering. So I’ve been thinking about the effect of splitting, too, and whether we need it. I asked Andrew and he said Kenny Shirley once produced an example where split R-hat diagnosed non-convergence that regular R-hat didn’t, but that example is lost to time and we’ve never seen this kind of behavior with NUTS as far as I know (please give us an example in the comments or via email to Andrew if you have).

Relating R-hat and ESS

My intuition was that we could set a low enough R-hat threshold that it would ensure a high enough effective sample size (ESS) when we crossed it. The relation’s a little tighter than I thought, with

    Rhat^2 ≈ 1 + M / ESS,

where M is the number of chains and ESS is the effective sample size of all chains combined. There’s a multivariate proof in Vats and Knudson, 2021, Revisitng the Gelman-Rubin diagnostic, Statistical Science, page 2 and section 5 for details, but it’s pretty straightforward to get the intuition when you reduce R-hat^2 to (N-1)/N + var(chain-means) / man(chain-variances) as Charles Margossian did in his nested R-hat paper. Vats and Knudson disapprove of Andrew and Aki’s suggested threshold of 1.1 from BDA3, because it is satisfied with a combined ESS of 20 across Andrew’s default 4 chains.

Being me, I tried to validate my intuition with simulations rather than linear algebra. Also, I like to see that things work in practice that theory entails to make sure I’ve understood all the assumptions baked into the theory (one can’t prove anything without assumptions!). When asked to code a simulation using ArviZ, Claude inserted a (2 * M) in the numerator in place of the M. Where did that come from, I asked? It told me it needed the factor of 2 because ArviZ uses split Rhat. D’oh! Of course it does, because we’ve doubled M without increasing ESS.

A worked example

Suppose we have 4 chains with a combined ESS of 400. Then sqrt(1 + 4/400) ≈ 1.005 and sqrt(1 + (2 * 4) / 400) ≈ 1.01. We’ve effectively doubled the number after the 1 by splitting. Unlike Vats and Knudson, I usually don’t need an ESS >> 100, so the 400 required for split R-hat < 1.01 is perhaps a bit too conservative for my tastes. On the other hand, we face a practical problem estimating ESS reliably with fewer than 50 or so ESS per chain. Estimation is challenging because it relies on autocorrelation estimates from the chains themselves, which become much noisier when based on shorter chains. (Side question: Do we not combine autocorrelation estimates across chains to reduce standard error because some chains might not be mixing?) Also, we know this algebra wasn't a coincidence of 4 chains and 400 draws. The Taylor expansion of sqrt(1 + x) is the convergent sequence

    sqrt(1 + x) = 1 + x/2 - x^2 / 8 + x^3 / 16 + ...

When x < 0.1, the first-order approximation, sqrt(1 + x) = 1 + x / 2, is good.

The bottom line for practitioners

We need around twice as many draws to get below a fixed threshold with split R-hat than with the original R-hat.

The optimizer’s curse

The above sketch shows a decision tree.

The circles are uncertainty nodes and the squares are decision nodes. Read the tree from left to right: to start, there is uncertainty of which of the strata i=1,…,I you will be in. In any given stratum, you will have to decide between options 1 and 2, and for each of these decision options there is uncertainty about the payoff.

The goals are:

(a) Conditional on the stratum, pick the best decision. This is the local decision problem.

(b) Averaging over the strata, evaluate the expected value of the tree, that is, the expected value under an optimal decision analysis given the uncertainty.

The challenge is that you don’t know which internal decision is best, because there is uncertainty about the payoffs.

The “optimizer’s curse” is that if, for each stratum in step (a), you make the best decision given available information–that is, you estimate the expected payoff under each of the two decision options and then pick the the one whose expected payoff is higher–then if you use these expected payoffs in step (b) you will systematically overestimate the value of the tree.

The “curse” here is not that the optimizer is making bad decisions, it’s that a naive estimate will be overly optimistic about the net value because you’re selecting on choices that look good.

In 2007, Erwann Rogard, Hao Lu, and I published a paper on the topic, including the above diagram. Here’s our abstract:

The evaluation of decision trees under uncertainty is difficult because of the required nested operations of maximizing and averaging. Pure maximizing (for deterministic decision trees) or pure averaging (for probability trees) are both relatively simple because the maximum of a maximum is a maximum, and the average of an average is an average. But when the two operators are mixed, no simplification is possible, and one must evaluate the maximization and averaging operations in a nested fashion, following the structure of the tree. Nested evaluation requires large sample sizes (for data collection) or long computation times (for simulations).

An alternative to full nested evaluation is to perform a random sample of evaluations and use statistical methods to perform inference about the entire tree. We show that the most natural estimate is biased and consider two alternatives: the parametric bootstrap and hierarchical Bayes inference. We explore the properties of these inferences through a simulation study.

I kinda like the paper. I wouldn’t say it’s one of my all-time favorites, but I think it’s interesting, and I like that we offer two different solutions to the problem.

On the downside, the paper seems to have disappeared without a trace. In 20 years, it’s only been cited three times, and none of them look very impressive:

“Using Alternating Decision Treets,” indeed.

Maybe one problem with our paper was its dry-as-dust title, “Evaluation of multilevel decision trees.”

This all came to mind because Sean Manning pointed me to this post, “The best cause will disappoint you: An intro to the optimisers curse.” Now that’s a good title.

It seems that the term “optimizer’s curse” came from this 2006 paper by James Smith and Robert Winkler, which has a lot of overlap with our article that appeared a year later. Both papers use hierarchical Bayesian analysis. Their paper is better than ours, for sure, and not just in the title, as they make a much better case for the importance of the problem. But we were working independently. Too bad: had we joined forces we could’ve produced something better, as each of the two papers had lots of material that was not in the other. Smith and Winkler consider the problem of choosing among many options with different levels of uncertainty, whereas we consider a multiplicity of binary decisions. These are just two cases of the general principle.

The above-linked post, by someone who goes by the handle “titotal,” is good too. It doesn’t have any new technical material, but it explains the problem in plain English from first principles, goes through some examples, and discusses some of the policy implications.

Bayesian Workflow exists as a physical book!

We’re very excited about this book. It’s the result of several years of effort. You can order from the publisher or from Amazon.

Here’s the book’s webpage, which includes the data and code for the book’s examples and case studies, of which there are many.

Here’s the table of contents:

Part 1: From Bayesian inference to Bayesian workflow
1. Bayesian theory and Bayesian practice
2. Statistical modeling and workflow
3. Computational tools
4. Introduction to workflow: Modeling performance on a multiple choice exam

Part 2: Statistical workflow
5. Building statistical models
6. Using simulations to capture uncertainty
7. Prediction, generalization, and causal inference
8. Visualizing and checking fitted models
9. Comparing and improving models
10. Statistical inference and scientific inference

Part 3: Computational workflow
11. Fitting statistical models
12. Diagnosing and fixing problems with fitting
13. Approximate algorithms and approximate models
14. Simulation-based calibration checking
15. Statistical modeling as software development

Part 4. Case studies
16. Coding a series of models: Simulated data of movie ratings
17. Prior specification for regression models: Reanalysis of a sleep study
18. Predictive model checking and comparison: Clinical trial
19. Building up to a hierarchical model: Coronavirus testing
20. Using a fitted model for decision analysis: Classification competition
21. Posterior predictive checking: Stochastic learning in dogs
22. Incremental development and testing: Black cat adoptions
23. Debugging a model: World Cup football
24. Leave-one-out cross validation model checking and comparison: Roaches
25. Model building and expansion: Golf putting
26. Model building with latent variables: Markov models for animal movement
27. Model building: Time-series decomposition for birthdays
28. Models for regression coefficients and variable selection: Student grades
29. Sampling problems with latent variables: No vehicles in the park
30. Challenge of multimodality: Differential equation for planetary motion
31. Simulation-based calibration checking in model development workflow

Appendices
A. Statistical and computational workflow for Bayesians and non-Bayesians
B. How to get the most out of Bayesian Data Analysis

One way to think of the book is that it’s all the things missing from BDA, like how to set up an informative prior, what to do when your computations aren’t converging, how to work through a series of models fit to the same data, how to design and perform simulated-data experiments . . . and all sorts of other things too.

The core of the book–parts 1 through 3–clock in under 200 pages, and then we have another 300 pages full of case studies demonstrating different aspects of Bayesian statistical and computational workflow. The appendices should be useful to you too, first because the workflow ideas in this book apply to non-Bayesian inference too, and second because BDA still has lots of valuable material in it, so it’s good to know where to look.

This new Bayesian Workflow book could change your life (we hope), and I thank my coauthors, Aki Vehtari and Richard McElreath, with Daniel Simpson, Charles C. Margossian, Yuling Yao, Lauren Kennedy, Jonah Gabry, Paul-Christian Bürkner, Martin Modrák, Vianey Leos Barajas, for all their care and effort. We thank our employers and various funding agencies for giving us the resources to be able to write this book as a side project along with all our daily responsibilities. And we thank many people for their input on earlier versions of the book, along with the Stan developers making so much of this work possible and the Stan community of users for supplying a continuing series of challenges that have motivated many of the ideas and methods discussed in the book.

I posted this already on the blog and you can see answers to some questions in the comments there. I’m posting it again here because, hey, we don’t come out with a new book every day!

I hope you find the book readable, interesting, and useful.

Structural equation modeling (SEM) and positive definiteness

This post is from Bob.

Mitzi and I were swotting up on structural equation models (SEM) for our class this past Monday at the Modern Modeling and Methods (M3) conference at Fordham University. It was a lot of fun and now I think I understand SEM notation. I really like these applied conferences and this was a group of psychometrician, econometricians, and sociometricians. Many if not most of them thought about models in terms of SEM, so we thought we should figure it out. But I was left with a concern you may be able to help me sort out.

The example

The first worked example in Ken Bollen’s seminal 1979 textbook on SEM is a study of how industrialization relates to democracy. It comes from his paper,

  • Bollen, Kenneth A. (1979). “Political Democracy and the Timing of Development.” American Sociological Review, 44(4).

and was reprised in his book

  • Bollen, Kenneth A. (1989). Structural Equations with Latent Variables. Wiley.

I had the pleasure of sitting across from Ken at the invited speakers dinner at the conference, so I’m glad I looked into SEM before that. Good news for the SEM devotees—he released a completely revised guide to SEM a few months ago.

The data and parameters

The data consists of eleven covariates (called “indicators” in SEM) for each of 75 countries. Four of the covariates are related to democracy in 1960 (y1, y2, y3, y4), the same four measurements were taken again again in 1965 (y5, y6, y7, y8) , and there were three measurements of industrialization in 1960 (x1, x2, x3).

The SEM model the original researcher came up with here assumes three latent scalars per country, industrialization in 1960 (IND60), level of democracy in 1960 (DEM60), and level of democracy in 1965 (DEM65). These latent parameters are related in the following way: democracy in 1960 is a regression on industrialization in 1960, and democracy in 1965 is a regression on both democracy in 1960 and industrialization in 1960.

The covariates are then modeled like a seemingly unrelated regression in econometrics. The four democracy 1965 parameters are treated as regressions on the latent level of democracy in 1965, and similarly for the democracy in 1960, and industrialization in 1960.

Rather than independent errors, a SEM model explicitly indicates with arrows which pairs of observations are allowed to have non-zero correlation in the covariance matrix for the observations. The three industrialization observations are assumed to have zero correlation—there are no arrows between any of the three measurements in the SEM diagram. Each of the four measurements in 1960 is assumed to covary with the same measurement taken in 1965. In addition, the second and fourth measurement in each year are assumed to be correlated with each other, which leads to a box-like structure.

The SEM diagram

Here are the arrows in the diagram, where I’m not using their standard LISREL notation, but writing them in R expression syntax to indicate what is regressed on what. In their graphical notation, just replace ~ with <-. All three latent variables and all eleven measurements are indexed by country.

IND60
DEM60 ~ IND60
DEM65 ~ DEM60, IND60

x1, x2, x3 ~ IND60
y1, y2, y3, y4 ~ DEM60
y5, y6, y7, y8 ~ DEM65

The covariance structure is indicated by stating which pairs of measurements are modeled with non-zero correlation. The first four just pair the measurements of the same thing across 1960 and 1965.

y1 <-> y5
y2 <-> y6
y3 <-> y7
y4 <-> y8

The last pair of correlations are within 1960 and within 1965.

y2 <-> y4
y6 <-> y8

Together, these induce an odd box structure, where y2 is correlated with y6 and y4, both of which are correlated with y8, but y2 and y8 are assumed to have zero correlation.

y2 <-> y6
^      ^
|      |
v      v
y4 <-> y8

Stan implementation

We didn’t get this far in my half of the class, so I will share here the Stan Playground example where I fit Bollen’s example (you can get the data and the Stan model through the Playground link:

It gets the right answer compared to lavaan/blavaan, which is nice. In the Stan code, xi is IND60 and eta1, eta2 are DEM60, DEM65. The relation among the latent parameters are modeled directly as regressions. The correlations among the observations are modeled using soft zeroing, where I just put a tight prior around zero on the structural zero elements, because Stan doesn’t give you a good way of setting up structural zeroes in a covariance matrix (Sean Pinkney or Ben Goodrich might know how to do this?).

This makes me curious how the lavaan package in R manages this. There’s a Bayesian version of lavaan built on top of Stan, blavaan. The first example right at the top of the home pages for both the lavaan and blavaan is Bollen’s democracy model. I guess it’s like the Scottish lip cancer data set for spatial modeling or Fisher’s iris data for regressions.

My questions

Consider a simple diagram among measurements like the following.

x <-> y
y <-> z

This says there can be non-zero correlation between A/B and also between B/C, but the correlation between A/C is zero. It’s a simplified case of the box we saw in the actual example. These arrows implies the correlation matrix looks as follows.

|        1  rho[x,y]         0 |
| rho[x,y]         1  rho[y,z] | = Omega
|        0  rho[y,z]         1 |

Given that the correlation matrix Omega must be positive definite, this limits the range of rho[x,y] and rho[y,z]. For example, we can’t have rho[x,y] = rho[y,z] = 0.9, or rho[x,z] would have to be greater than zero to maintain positive definiteness.

Q1: Why doesn’t SEM instead say that the correlation rho[x,z] is just the minimum value it can be given rho[x,y] and rho[y,z]? I’m suggesting that we instead treat the above diagram as implying no additional correlation between x and z other than that implied by the correlation between x and y and the correlation between y and z? That is, why try to shrink rho[x,z] all the way to zero? From the text, it feels like the motivation is to enforce zero correlation in the model. But all this is doing is simplifying regressions—it won’t actually enforce zero correlation among the measurements that are modeled with zero correlation. I wished I’d asked Ken this question at dinner, but I’ll ping him about this blog post and hopefully get a response.

Of course, in the pragmatic Bayesian workflow, we’d use posterior predictive checks to evaluate whether there’s unmodeled correlation between x and z.

Q2: I’m also curious what Andrew and others think about enforcing structural zeroes in correlation between measurements as opposed to just estimating a dense covariance matrix and inspecting where the correlations fall.