This University of Chicago business school professor has authored 258 academic papers in 2026 (so far).

Jeremy Horpedahl tells the story:

Nicholas Polson has, by my count using his SSRN page, already written 258 working papers in 2026 alone. He’s already written (or at least published to SSRN), six papers today, August 26, 2026.

OK, but who am I to talk?—I’ve written over 200 blog posts this year. But wait:

These aren’t just short notes. Most of the papers are of normal academic length: 32 pages, 27 pages, 58 pages. . . . Obviously the research productivity of Polson and his co-author Sokolov is aided by AI. . . . I have seen any academic, at least not in economics, that has really pushed it to the limit.

Horpedahl writes:

Read any single paper, and it feels like just a normal academic paper, the kind of thing that an academic might work on for a few months.

I wanted to see if I shared that judgment so I clicked through to the list of Polson’s papers on SSRN and looked for something interesting . . . ok, here’s something. It’s called Theories of Human Connection, and . . . ulp! It’s 80 pages long. The paper’s subtitle is “An Interdisciplinary Synthesis Across Economics, Psychology, Biology, Philosophy, Game Theory, and Spiritual Tradition.”

But let’s take a look. The abstract on SSRN starts like this:

Human connection is the most studied and least integrated phenomenon in the social sciences. Every discipline that examines intimate relationships captures something real that the others miss, yet no existing framework holds all dimensions in simultaneous view. This book synthesises fourteen thinkers across biology, psychology, economics, game theory, communication theory, existential philosophy, and spiritual tradition into a unified account. Morris established that the need for physical touch is an evolutionary drive as fundamental as hunger, with a biologically ordered sequence of escalating vulnerability whose disruption produces intensity without depth. Bowlby showed how early caregiving creates invisible templates governing adult intimacy, encoded in the nervous system before language exists. Becker revealed that partnerships generate value neither partner could produce alone, but assumed partners are interchangeable — an assumption Frankl demolishes by showing that the meaning generated by shared life cannot be transferred, and that meaning multiplies rather than adds to life satisfaction, explaining widespread disconnection in the wealthiest societies in history. Von Neumann’s game theory explains how mutual self-protective withdrawal, individually rational for each partner, produces the disconnection neither intended. Bateson identifies the communication structure that accelerates this collapse: contradictory demands that make any response wrong. Gottman’s laboratory models predict separation with over ninety percent accuracy from the ratio of positive to negative interactions. The Hindu philosophical tradition adds the final dimension: partnership as a laboratory in which selfishness and fear are progressively revealed and surrendered.

We argue that most relationship failures are not failures in one dimension but misidentifications of which dimension is actually in play, and that effective intervention requires the multi-dimensional map this synthesis provides.

From the preface:

The thinkers assembled here — Becker, Bateson, Von Neumann, Schelling, Keynes, Morris, Vaughan, Frankl, Maslow, Gottman, Yogananda, Vivekananda, Maharaj, Polson, Thomas, and Paltrow — did not, for the most part, know each other’s work. They worked in different centuries, different countries, different intellectual traditions. What unites them is that each identified a dimension of human connection that the others left in shadow.

Gottman, huh? The name rings a bell . . . He’s the guy who conned Malcolm Gladwell and various media outlets—maybe he conned himself too—into believing that he could predict divorces with 94% accuracy.

I searched for Gottman in the document and found a whole chapter on that bullshit! You can click through for yourself, it’s chapter 10. Here’s a key bit:

I wonder what prompts were used by Polson and his coauthor to write this article. I guess the prompts did not include, “Evaluate implausible claims skeptically.”

They bring it all together in Chapter 13, “The Cumulative Model: A Unified Theory of Human Connection”:


But let’s not forget “The Master Equation” on page 75:

Jesus Christ. I’ve heard that Cambridge University has an open position in their school of education . . . this kind of thing would fit in very well there, no?

In all seriousness, no, I don’t think this “feels like just a normal academic paper, the kind of thing that an academic might work on for a few months.” At least, not the sort of thing a non-bullshitting academic might write.

I know Nick Polson—he’s a statistician, and he’s done lots of solid work over the years! What happened here?

Here are a few possibilities:

1. It’s an experiment or a joke. But if it were a joke I’d think there’d be some internal clues, no? It’s hard to imagine playing the whole thing straight. And if it were an experiment, I’d expect they’d all read like straight-up statistics papers, nothing so obviously bogus as “Theories of Human Connection.”

2. Someone else is impersonating Polson. Seems unlikely, but it’s possible. It might not even be personal. Maybe all this is an experiment, not by Nick but by someone else who programmed a chatbot to choose the name of a successful academic and then spew the internet with papers attributed to him. If so, how horrible.

3. “Intellectual squatting.” That’s the conjecture of Michael Makowsky, who writes:

The nice version is it’s putting out a series of half-baked papers in the hopes of establishing a property right to the underlying ideas at an earlier stage of the research process than previously possible. The less generous interpretation is it’s dumping a series of haystacks on the plains and laying claim to the needles probabilistically within each.

I guess . . . but what does Nick ultimately get out of it? Invitations to speak at more conferences?? I don’t get it.

Makowsky writes:

Imagine you are a person who has highly esoteric, potentially important ideas every day. Many of those ideas you suspect, based on some combination of experience and ego, are new in at least one dimension. You would like to get credit for that newness. For being first. What’s the problem?

The problem is that scholarship remains more perspiration than inspiration. Having a new idea is great, but it takes years to work through the nuance in sufficient detail that you can convince your peers of the coherence and originality of the contribution. During the minutes each day you are not working on this singular project you have the inspiration for other ideas, sometimes multiple within a single day. How frustrating is the proposition that someone else gets credit for the originality of contribution just because they had time to reveal it to the world while you were embroiled in your investigation of what is only one of your many score ideas!?

Ah, but meta-level inspiration has struck you! What if you took each one of those ideas, spent an hour curating a series of prompts around it, and then let Chat GPT (or another LLM) fabricate an entire research paper around it? . . .

But . . . that’s what blogging’s all about! Often when I have an idea or a reaction, I blog it. No need to pipe it through a chatbot; I’ll just save the cycles and post it right here.

To return to the 256 working papers, I can think of one more motivation:

4. Education. This seems like the most plausible explanation to me. Polson has had an active research and teaching career, and he’d like to share his insights with a broader audience than the readers of his published papers and the students at the University of Chicago business school. And one way to reach people is . . . econ preprints! So Nick picks 258 interesting topics, writes some prompts for each, and produces the articles. I guess he’s programmed a bot to do this. He just feeds it the prompts and the bot writes the paper and posts it directly to SSRN.

That could explain the mystery of how that ridiculous 80-page article with “The Cumulative Model: A Unified Theory of Human Connection” (shades of Stephen Wolfram!) ended up there. Not only can’t you expect an author to write 258 articles of that length in less than a year, you can’t expect him to read all of them too. The content of that bizarre article could be as much a surprise to Polson as it was to me.

This then raises a question: setting aside the motivations of Polson (or his impersonator), do these 258 papers have any value?

It’s hard for me to answer this question, given that I’ve only looked at one of them. My guess is that the net value of the papers is negative, in that the amount of time that people (including me) have wasted going through them outweighs any positive contributions that might have been there.

My suggestion

Here’s what Nick could do on this, which could have value: Take these 258 prompts and write an article (himself, not using the chatbot) explaining why he thinks these ideas are important. Aki and I wrote a paper a few years ago, What are the most important statistical ideas of the past 50 years?. Nick could write something similar: What are the 258 most important things in statistics to learn today? Or something like that. I’m not saying it would be easy—it would take more effort than programming a chatbot to spam SSRN—but valuable products often take work to produce. Nick has tenure and could set aside the time to do it.

Also I’d recommend withdrawing all those papers from SSRN. Withdrawing 258 papers seems like a lot of work, but I’m sure he could easily program a bot to do the job.

P.S. There’s a further twist: there are two accounts for Nicholas or Nick Polson at the University of Chicago business school; see this comment thread. This would seem to be consistent with the “social experiment” hypothesis (if Nick decided to set up a separate account to play around with) or the “impersonation” hypothesis (if the bot that wrote and posted these papers was not created by Nick at all). The whole thing remains a mystery to me.

P.P.S. OK, I did a bit more nosing around.

SSRN allows you to list the papers in time order. If you go to Nick’s SSRN page linked from his website, you’ll see 16 papers, with the first (“The Impact of Jumps in Volatility and Returns”) being posted on 1 Jan 2001, then others through the next two decades, with the most recent being “Deep Learning in Characteristics-Sorted Factor Models,” posted on 23 Sep 2018 and last revised 26 Jun 2023.

If you go to the SSRN page with all the fake papers, it starts with “Kramnik vs Nakamura or Bayes vs p-value,” posted 7 Dec 2023. It’s a badly written paper—I’m guessing not AI, just text by a non-English-speaking author that was not ever checked by native speakers before posting. This rings a bell . . . I actually have a blog post on this paper, scheduled to appear next year. Next on the list is a 25-page paper, “AI and Vivekananda,” posted 5 Mar 2024, then a gap of two years until another AI-related paper appeared on 9 Mar 2026, then on 11 Mar 2026 came the aforementioned “Theories of Human Connection.”

So, yes, Polson has two SSRN pages, but they have no overlap in time. He also has papers on Arxiv, including the intriguingly-titled “Bayes with No Shame: Admissibility Geometries of Predictive Inference,” dated 24 Aug 2026 . . . Hey, that’s just 3 days ago! Oddly enough, I can’t find this one on SSRN.

But what about the article itself? I don’t have the patience to read it, but I did catch that it mentions the martingale property, which is one of my current interests—that’s cool. But, just flipping through, it looks much more substantive—much more like a real scientific paper—than that horrible “Theories of Human Connection” thing. This could be a tribute to the power of modern chatbots to create something so convincing.

P.P.P.S. Update here: The incredible shrinking SSRN page.

Postdoc and doctoral student positions in Bayesian workflow at Aalto, Finland

This job ad is by Aki

I’m looking for postdocs and doctoral students to work on Bayesian workflow. The candidates need to have knowledge of Bayesian inference and some experience with building models (for real applications, as part of methods development, or as part of courses). Although we have published Bayesian workflow book, there is still a lot more to do. The focus in the group is in cross-validation, model checking and inference diagnostics (see my publication list).

All positions are fully funded and the salaries at Aalto CS are 53k€-55k€ / year for postdocs and 40k€-45k€ / year for doctoral students. There are occupational healthcare and other benefits. Postdoc positions are typically offered for up to three years and doctoral student positions for four years. Starting dates are flexible and the details of each position will be agreed individually.

You can apply via joint ELLIS Institute Finland call and pick me as your preferred supervisor.

(1) “Do you think the culture of research has genuinely changed since the replication crisis became widely discussed, or has it mostly generated new compliance rituals around pre-registration and open data while leaving the underlying incentive structure intact?, (2) Regarding Columbia University, “is there a statistical or social scientific way of understanding how institutions lose the ability to accurately perceive their own situation?”

Luke Ford writes:

[Regarding] the replication crisis, researcher degrees of freedom, and the gap between what statistical methods claim to establish and what they can actually support . . . Looking at the current state of the social sciences, do you think the culture of research has genuinely changed since the replication crisis became widely discussed, or has it mostly generated new compliance rituals around pre-registration and open data while leaving the underlying incentive structure intact?

And a question about Columbia specifically since you are there: the university has had a difficult two years in ways that have played out publicly. From your position as someone who thinks carefully about institutional incentives and measurement, what do you think the administration consistently misread, and is there a statistical or social scientific way of understanding how institutions lose the ability to accurately perceive their own situation?

My reply:

1. I’m loath to give an answer about the changes in the culture of research because I have not studied this systematically. My impression is that, yes, there’s more skepticism and less acceptance of noisy N=38 papers in psychology, etc., and less toleration for unfalsifiable evolutionary psychology and that sort of thing. On the other hand, perhaps this has just shifted from the science establishment to social media. Ten or fifteen years ago, there was a pipeline (partly abetted by Jeffrey Epstein) from researchers at top universities to publication in top journals to books, NPR, Ted, Gladwell, Freakonomics, etc., and lucrative speaking and consulting gigs. So you get people like Marc Hauser or Albert-Laszlo Barabasi or Brian Wansink or Dan Ariely doing the basic research (such as it is), academic middlemen such as Steven Levitt and Cass Sunstein as promulgators, and the universities, journals, and prestige news media as part of this system (as for example here: https://statmodeling.stat.columbia.edu/2023/08/31/the-variation-ignoring-junk-science-thats-promoted-by-association-for-psychological-science-and-related-academic-celebrities-its-like-a-poker-player-thinking-okay-if-push-all/).

Nowadays, though, social media runs on its own steam, and the models for academic junk science are researchers such as Andrew Huberman and Dr. Oz, who cut out the middleman and promote junk science directly, sell supplements, etc. And social media is full of fake news and AI slop. They don’t really need NPR, Ted, Gladwell, Freakonomics, etc., anymore; they can do it on their own. So, in short, yes, I do have the impression that science has reformed from the bad old days of 2010-2015 (about which, see this article with Simine Vazire: https://sites.stat.columbia.edu/gelman/research/published/jmmss-3062-gelman.pdf), but maybe the public intellectuals don’t need academic science anymore; they can just make up whatever they want on their own.

2. My take on Columbia is similar to my take on many institutions, which is that they have an executive function but minimal legislative or judicial functions; I discussed this here: https://statmodeling.stat.columbia.edu/2018/01/19/lesson-charles-armstrong-plagiarism-scandal-separation-judicial-executive-functions/ and here: https://statmodeling.stat.columbia.edu/2025/11/11/from-the-three-branches-of-government-to-the-bidirectional-nature-of-legal-reasoning-in-a-way-that-is-similar-to-how-statistics-works-and-should-work-in-the-real-world/. As a result, their decisions are made on consequentialist rather than proceduralist gounds, and over and over again the administration takes the seemingly reasonable decision to cover up misdeeds.

Ford posted this discussion on his blog. It was kinda weird seeing myself discussed as a sociological object, but, fair enough, I’m a public figure, and people can say what they want as long as they don’t misrepresent my writings or claim that I said something I never said.

Ford’s assessment is accurate that I’m not very good at strategic behavior so often I don’t even try. It’s similar to how I’m a bad negotiator so usually I’ll just try to make my goals clear and not try to optimize, following the “Getting to Yes” principle that the main thing getting in the way of smooth negotiation is ignorance of other people’s goals. I think back to various successful and botched negotiations I’ve been involved with in the past, and almost always the problems come with struggles over details without there being clarity on the goals of the different parties.

Survey Statistics: more on SynthMargins and Bayes-Raking

Last week we discussed SynthMargins, a method from the poster Modeling Complex Contingency Tables that uses partial information (margins) about poststratification variables. On theme for this series (“it is the people”), the authors commented thanking Andrew for introducing them, folks from different fields with a shared goal: Max GoplerudShiro Kuriwaki, Jens Wiederspohn, Adam Conner-Sax, and Philip Greengard.

Bob Carpenter shared a Stan example to get a flat prior over tables that match specified margins. I think this prior would imply a prior on what the authors call alpha0, the covariance coefficients for the target geography. The method as described in the Modeling Complex Contingency Tables poster seems to focus on point estimates:

In other news, Shiro responded on Twitter to one of my questions about their Application 1 (ACS): “the density plot shown there is an empirical density of 1700+ TREs, where each error is for a non-Southern county.”

I have 2 remaining questions:

  1. What are the covariates w in this ACS example ?
  2. Suppose you also have survey data in Palo Alto. Would this be added to the training tables ?

Their Application 2 asks if lower postratification table reconstruction error improves the downstream MRP:

In the comments last week, Shiro cited related work by Si and Zhou (2021) who propose a method called Bayes-Raking to incorporate known margins into modeling. They found Bayes-Raking was similar to raking in the overall mean but outperformed raking for subgroups:

It would be interesting to directly compare Bayes-Raking to the poster’s SynthMargins !

Bayesian Workflow free pdf!

Our wonderful new Bayesian Workflow book is now available as a free pdf! Just go the link—it’s right there!

I recommend getting the hard copy too because you’ll want to be able to read it while working on the computer, and the cost of the book is trivial compared to the benefit from faster learning that you will get by being able look at the book without taking up valuable screen real estate; also you can see connections when flipping through the pages that might not be apparent by viewing one page at a time on a screen.

Conversely, if you have the hard copy, you should still download the pdf because it fixes a bunch of minor errors that we caught after the book went to press. Also in the printed version we accidentally repeated some of the exercises in chapters 2 and 3. For the pdf we fixed this.

Regarding the content, as I wrote last month, with Bayesian Data Analysis, the big steps forward were:

  • Going beyond Bayesian inference to also consider Bayesian model building (as a researcher, you construct the model, it isn’t just given to you as in a textbook), model checking (breaking through the absolutely horrible attitude, common to Bayesians in the early 1990s, that the model was “subjective” and thus should not be checked), and model improvement (continuous model expansion, not the misguided idea of assigning posterior probabilities).
  • Going beyond simple conjugate models. BDA had lots of hierarchical models, also lots of computational tools so that you could fit the models you want by putting them together from understandable components. And I like how we had a clear separation between modeling and computing. The model comes first, then you figure out how to compute it. Or you set up a model that works within your computational constraints.
  • A Bayesian approach to sampling and causal inference. This was Rubin’s framework in which unobserved units in the population and unobserved causal outcomes are treated as missing data and are part of a joint probability model. We worked this out in chapter 7 of BDA (which became chapter 8 in the third edition of the book).
  • Lots of live examples. Not just “real-data examples,” but problems we’d directly worked on. This motivated us and I think it gave our readers a sense of how Bayesian methods worked not just in theory but in applied problems.
  • A pragmatic view of probability as a measurable quantity. That’s right there in chapter 1. Bayesian methods are not the product of a philosophical stance; they’re a way to connect models and data using probability.

I could go on and on, but for that I can refer you to the Bayesian Data Analysis book.

And these are the key innovations of Bayesian Workflow:

  • Going beyond Bayesian data analysis (model building, inference, model checking, and model expansion) to consider the larger process of statistical modeling, including comparisons of multiple models fit to a single dataset.
  • A fuller use of informative priors. This is a big deal. In BDA we still had a bit of the Bayesian cringe going on. One reason we’ve moved toward stronger priors is that the replication crisis has taught us that the amount of prior information available in any given problem is often approximately the same as the information coming from an experiment (see here, for example). Informative priors also fit our increased focus on generative modeling, and we’re doing a lot more prior predictive checking to understand the implications of our models.
  • More integration between modeling, data analysis, and computing. One way to see this is that the Bayesian Workflow webpage has the code to run all our examples. We also have lots of code snippets in the text as a way of demonstrating the way in which coding is central to our statistical workflow.
  • Lots more live examples. It’s been 30 years since BDA first came out. One reason that Bayesian Workflow has 11 authors is that different collaborators worked on different examples (but the three principal authors read through the entire book, so the general approach should remain coherent).
  • Simulation-based experimentation. This is something my colleagues have been doing more and more over the years. At its most basic, simulation-based experimentation provides a best-case baseline for statistical methods: if you can’t recover your quantities of interest with sufficient accuracy under ideal conditions (when your data are simulated from the model you’re fitting), then you know you’re in trouble. And often this is the case! Beyond that, we can simulate from one model and fit another, and see what happens. Simulation experiments aren’t always so easy to construct, as they involve specifying the entire data-generation process. But we think this is effort worth expending, as it involves thinking about the problem you’re working on.

I could go on and on, but for that I can refer you to the Bayesian Workflow book.

What contributions can academic statisticians make to sports analytics? (a discussion related to the launch of the new open-access Journal of Statistics and Data Science in Sports)

Related to our recent post on the absurdity of open access fees, somebody recently informed me that a group of statisticians that work in sports decided they were sick and tired of this with the Journal of Quantitative Analysis in Sports and launched a new open access journal, the Journal of Statistics and Data Science in Sports:

JSDSS was founded on three core principles.

First, our commitment to open access is more than just lip service. At JSDSS, open access means free to read and free to publish—no exceptions. JSDSS is a Diamond Open Access journal. Research supported by academic institutions, public funding, or personal effort should not require a payment to reach the audience it deserves.

Second, reproducibility is essential and not an afterthought. The credibility of sports analytics research depends on our ability to verify, replicate, and build upon earlier work. JSDSS will actively incentivize transparency in data, code, and methodology, and work that meets our reproducibility standards will receive a special designation recognizing this commitment.

Third, we believe sport is a rich and underutilized laboratory for statistical and data science innovation. From player evaluation and in-game strategy to league design and fan engagement, sports data present compelling, real-world problems that demand rigorous and creative analytical thinking. We intend for JSDSS to be the definitive venue for this work.

Cool! I should send them something. We think about statistics and data science in sports a lot around here.

My only concern is that I get the impression that the cutting-edge work on sports analytics is happening outside of academia. So I hope that this new journal can get useful contributions from people who are working in sports analytics who have material they can share without compromising their competitive advantage.

What contributions can academic statisticians make to sports analytics?

Or we could flip it around and ask, What contributions can academic statisticians (like me!) make to sports analytics? Here are a few things:

– Developing general methods that can then be used in sports analytics (as here);

– Writing textbooks explaining general methods that can then be used in sports analytics (as here);

– Teaching students who can then work in sports analytics, or consulting on sports analytics projects, which can be thought of as a form of intense teaching;

– Doing work in sports analytics which, although it is not cutting edge, can still give valuable insights (as here);

– Evaluation and criticisms of published work in sports-related topics (as here);

– Contributing to the sports analytics community (as here, and indeed as in the present post);

– Collaboration on sports strategy, or sports medicine, or the sociology of sports, or various other places where sports links up with academic research;

– Clearing up confusion on topics related to statistics and sports (as here), along with social-sciency thinking about how these misconceptions persist;

– Sports-related research where it can be helpful to have an outside perspective, something available to academics who aren’t on a deadline (as here).

There are probably some more things I didn’t think to include on this list.

Head to head on 125 St: The Jamaican beef patty battle!

My new posts here have a one-year waiting list, but I’m bumping this one up because it’s important.

OK, we went on over and did a head-to-head Jamaican beef patty taste-off. As the above wrappers indicate, we compared the spicy beef.

None of this randomization, tea-tasting crap, we just kept taking bites of each patty until they were done.

Both were good, but the classic Golden Krust was definitely better than the newcomer, Juici Patties. The Golden Krust patty had a more delicious filling which was also better integrated with the crust. By comparison, the Juici Patty had a bit too much crust and was too empty inside—it didn’t work as well as a whole. Also the Golden Krust patty was $4.05 and the Juici was $4.25.

In summary:

1 Golden Krust patty > 1 Juici patty >>>> 1/1381 of a conference featuring Grover Norquist, Gray Davis, and a rabbi.

Failing upward, Norwegian style

Wendy Moore, in a review of a book by Oliver Basciano, writes:

Leprosy is caused by a bacterium, Mycobacterium leprae, first identified by a Norwegian doctor, Gerhard Armauer Hansen, in 1873 . . . Convinced, wrongly, that leprosy was hereditary, he spurred Norway to introduce the Seclusion of Lepers Act (1885), which enabled the authorities to remove people from their families to isolated leprosaria.

OK, fine, everybody makes mistakes, better to err of the side of caution bla bla blah.

But then comes this stunner:

After Hansen injected a virulent strain of the disease into the eye of one patient, Kari Nielsdatter Spidsøen, without her consent, she took him to court. He was stripped of his hospital post in 1880, but continued to oversee Norway’s leprosy policy.

Whaaaa?

Further research (i.e., I went to the Hansen’s wikipedia page) yielded this:

Hansen had attempted to infect at least one female patient with the nodular form of leprosy without consent, and although no damage was caused, the case ended up in court and Hansen lost his post at the hospital.

“No damage was caused,” huh? He just injected her in the eye, that’s all. No harm, no foul, I guess.

Just amazing that he continued to run government policy after that. Kinda reminds me of how noted terrorist John Poindexter was tasked by the U.S. government to run a terrorism prediction market. I guess it makes sense—he was a true expert on the topic.

Also reminds me of that unfortunate psychology researcher who keeps coauthoring fraudulent research papers, got disciplined by MIT, and then left for a prestigious chair at Duke University, ran an advice column in a major newspaper, and had a TV show made about his research. Actually two TV shows: one is a highly critical documentary and another is a fictional show where a character based on him is the hero.

Or that political figure, I can’t remember his name now, who keeps citing discredited, fraudulent, fake, and racist research claims, and at the time of this writing he remains in charge of health policy in a major industrialized country.

Some people end up in prison for minor crimes. Other people do things like inject people in the eye with a virulent strain of leprosy, and they get to make government policy.

This is one of the worst scientific papers I’ve ever seen.

A frequent commenter pointed me to this paper, “Sport and longevity: an observational study of international athletes.”

All I can say is . . . Wow! This paper is an absolute clinic in bad quantitative social science research.

I’ll leave it as an exercise for the reader to count up all the problems.

The key takeaway: For God’s sake don’t play volleyball. It’ll reduce your life span by 5 years, and the result is statistically significant, with a p-value of 2.4e-11.

I used to play some volleyball. Our team made the B-league playoffs in the MIT intramural league one year. Now I’m scared. Who knows what all that jumping did to our fragile young bodies.

Also this:

Your tax dollars at work!

Too bad the journal doesn’t seem to publish its reviews. It would be hilarious to see the referee reports for this one.

P.S. To those of you who think I’m being mean here, “punching down,” etc., let me just say a few things.

1. It’s not personal. I know nothing about the authors of this paper and I purposely did not include their names in the post. They may be wonderful people, indeed they may do wonderful research in other areas.

2. As noted above, public funds were spent on this project. And the paper was published, i.e. it’s there for anyone to read. If you don’t want to be criticized, don’t publish.

3. Unsupported claims get out there and they don’t go away. For example here’s what happened when I googled *pole vaulting longevity*:

This is just one silly example. But, yeah, I do think we should be bothered by broadcasts of unsupported scientific claims, whether they’re about volleyball, himmicanes, faith healing, air rage, governors’ lifespans, or anything else. It’s bad science and it’s part of our culture of B.S. Even if the authors of this particular paper are completely sincere in their efforts. Remember, honesty and transparency are not enough.

Here are the talks from StanCon 2026!

StanCon 2026 just happened!

And here are the talks:

Matthew Kay
Adrian Seyboldt
Paul-Christian Bürkner
Charles Margossian
Javier Enrique Aguilar
Sean Pinkney
Nikolas Siccha
Anna Dreber
Kaitlyn Johnson
Jonas Wallin
Pranav Sanke
Anna Elisabeth Riha
Colling Cademartori
Tim M. Szweczyk
Bob Carpenter
Chandler Ross
Nils Rudi
Fredrik Ronquist
Jakob Torgander
Aleksi Lahtinen
Ville Laitenen
Zeno Romero
Soham Mukherjee
Steve Bronder and Brian Ward
John Ashley Burgoyne

Lots of great stuff here. Check out the titles of the talks—all sorts of different topics.

Time to start planning for StanCon 2027.

My answer is No.

It would be great to have 35% more citations and over 5 times as many downloads, but I’d rather have 2,740 Jamaican beef patties.

P.S. Here’s the article in question: The ladder of abstraction in statistical graphics. I absolutely love this paper. It’s based on an idea I’ve had for awhile that I spoke on a few years ago in Ron Yurko’s statistical graphics class at CMU. Then I wrote it up and submitted it to the journal, which gave some useful comments, and I enlisted Kaiser Fung to get it over the finish line. Enjoy.

One night in Uzbekistan: Why was this one data point so influential, and what should these researchers had done ahead of time to see this?

Sol Hsiang writes:

We have a comment coming out in Nature next week that is going to cause the retraction of a high-profile paper by Kotz et al. from last year (the second most cited climate paper in the news in 2024).

Basically, Kotz et al claimed that climate change was already costing the world economy a huge amount and would cost 300% of what prior estimates claimed (which was already large). This result got enormous attention in Europe, in particular.

We couldn’t reproduce their findings and realized that it was all driven by weird data from Uzbekistan. If you remove Uzbekistan from their data set, the result falls apart. The costs are still large, but not the extreme numbers that made headlines around the world.

One reason we think this is important is because these data were previously being used by central banks around the world to run stress tests for the effects of climate change.

Here’s the retraction note, in full:

The authors have retracted this paper for the following reasons: post-publication, the results were found to be sensitive to the removal of one country, Uzbekistan, where inaccuracies were noted in the underlying economic data for the period 1995–1999. Furthermore, spatial auto-correlation was argued to be relevant for the uncertainty ranges. The authors corrected the data from Uzbekistan for 1995–1999 and controlled for data source transitions and higher-order trends as present in the Uzbekistan data. They also accounted for spatial auto-correlation. These changes led to discrepancies in the estimates for climate damages by mid-century, with an increased uncertainty range (from 11–29% to 6–31%) and a lower probability of damages diverging across emission scenarios by 2050 (from 99% to 90%).

The authors acknowledge that these changes are too substantial for a correction, leading to the retraction of the paper. An updated version of the paper with these changes, which has yet to undergo peer review, is publicly available with continued open access to its data and methodology (https://doi.org/10.5281/zenodo.15984134). The authors intend to submit a revised version of the paper for peer review. If and when published, this retraction note will be updated to include a link to the new publication. The authors appreciate the corrective role of the global scientific community and thank Thomas Bearpark, Dylan Hogan, Solomon Hsiang and Christof Schötz for bringing these issues to their attention. All authors agree to this retraction.

Good for them. And here’s the story in Retraction Watch.

How science advances when data and methods are open

Jonathan Falk independently pointed me to this story and wrote:

Imagine how uphill it would have been without access to the original data/methods.

Good point!

He also pointed to this news article which summarized the story:

If Uzbekistan were excluded . . . the damages would look similar to earlier research. Instead of a 62 percent decline in economic output by 2100 in a world where carbon emissions continue unabated, global output would be reduced by 23 percent. . . .

Wait—the estimate declines by almost a factor of 3 after removing just one data point? Uzbekistan’s not a tiny country but it’s not huge either (population 40 million); it doesn’t seem like its data should have so much influence as all that.

I went back to the original paper and it has some scatterplots, but (a) it’s hard to see that any one point would be so influential, and (b) the countries aren’t labeled so I don’t see which one is Uzbekistan.

A question of influence

What happened with the data? Is there some sort of scatterplot that would’ve indicated a concern?

To put it another way, if the data from a single medium-sized country could have that much of an impact on the findings, that would’ve been worth reporting from the get-go in the original paper. Even had there not been any data problems, we’d want to know that the results were so sensitive to one data point.

So the meta-question is: What data analysis should’ve been done originally, either to flag the problem with Uzbekistan’s data, or at least to reveal the extreme sensitivity of the headline results to that one data point?

I posed this question to Hsiang, who responded:

We noticed this issue because we were looking at several papers and running some basic diagnostics on all of them. One thing we were doing was just dropping one country at a time and rerunning the models to make sure things weren’t being driving by a single country. We were surprised that this turned up. There are many issues with this paper conceptually, but it’s not even really possible to discuss any of them until you deal with the UZB issue. We had a lot of dialogue with the authors, and it turned out that they really hadn’t run much quality control on the more granular data. When we traced back this issue, it seemed like their research assistants had faithfully converted some numbers from a PDF document, but those numbers were just implausible.

There is a scatterplot in their data paper that is supposed to provide technical validation of their data set. We wanted to see why Uzbekistan didn’t jump out, so we reproduced it in our comment (Extended Data Fig 1). It turns out that Uzbekistan wasn’t even the biggest outlier, but that the version they had published had the axes cropped so you couldn’t see the outliers (see red boxes in our version). This seemed indicative of a different issue, which is why we documented it in the comment.

I guess those graphs should be on the log scale?

It still seems crazy that the data from a single mid-sized country could have such a big effect of a global estimate. That’s something that the original researchers should’ve been aware of, and what it suggests to me is that there is a larger methodological problem that this didn’t get looked at automatically during the research process.

Why I am so against volunteering in academia (ecology-edition)

This post is by Lizzie. The photo is from this summer and included for no other reason than because I like a photo with a post. 

A few months ago I wrote a post where I compared the toxic culture of the restaurant industry to the toxic culture of some parts of my academic world in ecology. One point of similarity was how you often need to ‘volunteer’ your time to get a foot in the prestigious door. This led to a query by Phil about why I am against volunteering. Thanks to Phil for asking the query, and others for already effectively giving my general answer, but I will write a full post here since I think it is a good question.

So, why I am so against volunteering in academia? And here I am focusing on my part of academia, which is ecology. In short: I find volunteering in ecology in my world to be an exclusionary practice practiced by those who talk endlessly about inclusion. And, as is often the case in my life, this sort of hypocrisy drives me nuts and I can just never get over it.

And for the longer version….

In my field (ecology) and related fields (evolution, conservation), there is a pervasive assumption that it is a-okay to hire/have ‘volunteers’ to get your work done. I say hire, because these positions are advertised all the time with a list of qualifications you’ll need. Here’s one:

  • BSc degree in a wildlife, environmental or conservation topic or in the process of completing one.
  • Intermediate level in English and Spanish (Oral and Written).
  • Knowledge in wildlife monitoring surveys. Previous research experience in any capacity is a plus.
  • Physically fit and able to work long hours in a difficult and harsh environment.
  • Good team member with excellent communication skills; able to live and work with a multicultural team.
  • Able to live in basic living conditions and tropical rainforest conservation campus.
  • Hard working and passionate with a desire to learn and improve; willing to put the hard work in to go the extra mile for conservation efforts and personal career development.
  • Excellent computer skills with a confidence in all Microsoft and Google programs.

Conservation organizations or conservation-related research positions seem to feel especially allowed to do this, with the argument that their mission somehow absolves them of basic labor laws. The ones that maybe annoy me the most are the ones for ornithology work (that’s a fancy word for the study of birds) where you also need to pay your way to some (often tropical) locale and maybe even your housing and food so you can help slosh around in the jungle and search for birds. If you don’t believe me, most days you can set the Job Type to Volunteer/Training and search for ‘bird’ here and you’ll likely find something like this somewhere:

This is an unpaid position and interns are responsible to cover food and accommodation at the biological station [in Peru/Belize/etc.].

I learned about all this when I was on a ‘field crew’ long ago (that’s a term in ecology for a lot of people working together on some project where they go outside every day to collect data from the natural world). I was doing my PhD on bugs, but everyone else was studying the birds that eat the bugs and they explained to me that usually to get a paying bird-crew job you must first pay your way (housing etc. wherever the field crew is based) to be trained in point counts (standing in one place for a set amount of time and listing all the birds you hear) and then maybe someone pays your room and board and you learn to do ‘nest-finding’ (self-evident definition) and, as you get more and more training, you might get a small allowance (it’s an ‘allowance’ because it is generally way below minimum wage if you do out the per hour rate).

And then maybe you get into grad school to study birds. Lucky you!

And who knows how marine biology works (“Junior scientists in marine mammalogy are expected to have at least one or two unpaid research experiences to qualify for a graduate programme” according to a Nature Careers piece in 2020). Effectively, as best I can tell, the more coveted and sexy the position, the more it depends on extreme levels of unpaid work. I have never seen the data but when I look around at my ornithology and marine biology colleagues I wonder if the t-test on their socioeconomic status before they entered the ornithology/marine biology is higher than other subdisciplines within ecology.

And then the exact same disciples publish articles about how we need more diversity in the sciences. It blows my mind.

For decades people have been pointing out the problem:

Whitaker 2003 The Use of Full-Time Volunteers and Interns by Natural-Resource Professionals (Conservation Biology, Vol. 17, No. 1) explained that most often

full-time [unpaid] positions are filled by aspiring natural-resource professionals (e.g., students or recent graduates) who need workplace experience if they are to advance to graduate school or more lucrative jobs. Consequently, employers often consider work experience adequate compensation for wage shortfalls.

… and continues:

I believe that our widespread use of volunteers and interns to compensate for budget shortfalls does our profession more harm than good. In many cases this approach is in conflict with labor law, hinders the development of new professionals, undermines our profession’s credibility, and is an impediment to achieving our conservation goals.

He then argues that they these positions cause personal hardship, exclude certain groups, and does not meet societal norms:

Our perception of the importance and urgency of our work as conservationists does not elevate us above the societal values that led to [labor laws].’), undervalues conservation and those working in this and related areas.

Our current strategy of cutting wages does not cut costs; rather, it transfers them directly to the lowest tier of professionals in the form of financial and emotional hardship.

Despite this article from 20+ years ago my discipline continues to do this, now at the same time they decry the need for greater diversity in the field. Most articles on how to increase diversity preach for wider acceptance, statements of inclusivity, but not much action on this practice. That said, some articles acknowledge this as a problem we should maybe change if we want to increase diversity. They all cite Fournier & Bond 2015 Volunteer Field Technicians Are Bad for Wildlife Ecology (Wildlife Society Bulletin 39(4):819–821, which is good but I think the fact that they don’t cite Whitaker means they did not try very hard to look into this issue), which to me gives the 3 reasons folks usually give for not paying people who work for them:

  1. I don’t have enough money to pay them.
  2. It’s what has always happened and …
  3. pointing out other worse-treated/compensated workers.

Fournier & Bond 2015 also makes the general argument of why we should not have unpaid interns, with the Biden quote:

Don’t tell me what you value. Show me your budget, and I’ll tell you what you value.

I find (1) pretty common, but (2) is also pervasive. I have tried talking to my colleagues who are very well funded about this and they tell me that this as just ‘how science is done.’ They often explain to me that a student gets a letter of reference from them out of it, so what’s the problem?

I think once it’s engrained that you can get free grunt work, it’s hard to get people to give it up.

And folks are trained in it early. Graduate students are often sent out to get others to help with their grunt work. I see posts on ECOLOG and flyers in my hallway each year explaining who will be needed for a full-time unpaid position to collect soils, or leaf samples or pipette all day long.

What do I think everyone should do? I suggest they start by what I have managed to do: I just don’t allow ‘volunteers’ in my lab and I tell everyone why I don’t allow volunteers. I pay everyone who works in my lab (including undergraduates) and I make sure part of their paid word is training (how to use Git R, sometimes Stan). Is it expensive to pay every undergrad in my lab minimum wage (or better) given my super low budget each year? Hell yes. But that is not an excuse to me — if you cannot afford to get the work done without volunteers, then you cannot afford to get the work done.

Everyone is clearly not going to just do this though, so we could all benefit from some top-down help I suspect. Granting agencies could ask about how much a lab relies on volunteers and discourage it through various mechanisms. Instead, they often encourage it. In Canada, NSERC grades me on how many students I have ‘trained’ (AKA: have passed through my lab) so finding undergraduate volunteers to stock my lab would certainly help the numbers game. Further, NSERC cares about how much you have volunteered in MSc/PhD fellowships, effectively encouraging students to get on board with this practice.

I think NSERC means to support short-term volunteering, which gets back to Phil’s question of ‘when is this okay?’ In part he seems to ask if it is okay to take a minimum wage job on a rec crew or such. To which I say, of course! We’re all welcome to take low-paying jobs if we can get them if you ask me, especially as they are often the gateway to a career change and I support career changes. So the question is more: when is okay to work for free? I work on community-science projects, such as the USA-NPN, where data are entirely collected by volunteers and I see how volunteering in this way or in other ways that help you connect with something beyond your job, give back to community or provide aid are important for a multitude of reasons.

So my focus here is really on full-time or similar positions. Whitaker 2003 makes the distinction between ‘part-time or short-term service’ where you do not have to forego outside opportunities for paid work vs. “other volunteer and internship positions require that individuals live in remote areas and work>=40 hours per week, effectively denying them the opportunity to earn outside wages” and I think this is very good perspective — as it can be extended also cover the undergraduates working in my lab 10 hours/week during term (they do not forego getting part-time wages, which many students need to get through university).

But we should still be thoughtful of the diversity problems volunteering often leads to. Community scientists are often retired and in a good socioeconomic position, and this extends through many other places. If the National Parks Service relies on volunteers to help repair trails etc., I am sure they will have a bias in who signs up for those opportunities and that may not be the best thing for the NPS.

And while I am back on my diversity angle, if you are thinking it is okay for granting agencies like NSERC to ask about volunteering with something along the lines of ‘but don’t worry! If your life is too hard to volunteer, you can just explain that now and we’ll value it,’ then try harder.

Try harder to think through what type of person is the one who will know they just need to profligate themselves and tell all their horrible life problems so the nice middle- and upper-class people on the committee can give them credit for their difficulties. Maybe you get a small slice of people who that works for, but more I think you get several predictable outcomes. You get people lying about it. And you get people who are horrified and scarred by having to do it (IMHO it’s a revolting ask if you back up and think about), or just don’t do it. So while I am on the topic of suggesting we all stop asking people to volunteer to do our work, we should also stop asking people to tell us how hard their lives were.

What does it mean when different articles published by different people on the same topic have nearly identical titles? (“A novel nutritional supplement reduces postprandial glucose response in healthy individuals in a randomised, placebo-controlled, crossover clinical study”)

Sanjeev Sripathi writes:

I came across a firm that seems to be offering a miracle product – https://letsmoderate.com/products/sugar-slayer . Their claim is that ingesting it hammers your post-meal blood sugar spike down by 40%, and they have multiple studies to back it up: https://letsmoderate.com/pages/science

What stood out to me as weird:

(1) These three articles were published at different times by different people but follow the same format for the abstract and text and arrive at nearly the same very large effect size for the same product (the one being sold): https://www.mdpi.com/2072-6643/16/14/2237 , https://link.springer.com/article/10.1186/s41110-024-00275-6 , https://link.springer.com/article/10.1186/s41110-024-00294-3

(2) The other listed articles seem to have been fetched by doing a search for ‘mulberry extract’, which is the critical active ingredient. Whatever I can access does show a notable attenuation on blood glucose and studies exist all the way from 2007 to 2022, thought the impact varies a lot.

(3) These 3 studies are eerily similar to the first 3 I’ve mentioned here: https://www.liebertpub.com/doi/10.1089/jmf.2014.3160, https://link.springer.com/article/10.1186/s12986-021-00571-2, https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0172239#authcontrib . But they’re from 2007, 2015 and 2024, across 3 sets of researchers, which just seems very odd. Is this how everyone is meant to name their papers? i.e. is there a social pressure sitting outside these groups pushing them into common convention?

I’m sure you’d identify more elements if you took a look at it. I’m just unclear if I’m jumping at shadows.

I was curious so I clicked on the link. From the letsmoderate site:

Here are the titles of the first three articles listed in the above email:

A Novel Nutraceutical Supplement Lowers Postprandial Glucose and Insulin Levels upon a Carbohydrate-Rich Meal or Sucrose Drink Intake in Healthy Individuals—A Randomized, Placebo-Controlled, Crossover Feeding Study

A novel nutritional supplement reduces postprandial glucose response in healthy individuals in a randomised, placebo-controlled, crossover clinical study

A ready-to-mix nutraceutical supplement (GLUBLOC) lowers postprandial blood glucose levels in healthy individuals — A randomised, placebo-controlled, crossover study

All the above authors are from India but with no overlap in the author list.

And here are the titles for the other three:

Mulberry Leaf Extract Improves Postprandial Glucose Response in Prediabetic Subjects: A Randomized, Double-Blind Placebo-Controlled Trial

Mulberry leaf extract improves glycaemic response and insulaemic response to sucrose in healthy subjects: results of a randomized, double blind, placebo-controlled study

Mulberry-extract improves glucose tolerance and decreases insulin concentrations in normoglycaemic adults: Results of a randomised double-blind placebo-controlled study

The first one’s from Korea, and the second and third are from England, with some overlap in the author list.

I agree that it’s weird that the papers have nearly identical trials. But I don’t know how things go in the medical literature. Maybe this is standard practice?

I’m not planning to fork over 799 rupees for 30 tablets of Sugar Slayer. One of the article says, “Alkaloid- and polyphenol-rich white mulberry leaf and apple peel extracts have been shown to have potential glucose-lowering effects, benefitting the control of postprandial blood glucose levels,” so maybe I’ll just eat an apple.

Update on a regression discontinuity dispute: some asynchronous collaboration

Anjali Thomas writes:

I am writing to share a paper which is a re-examination of my 2018 AJPS article entitled “Targeting Ordinary Voters or Political Elites”  which was previously discussed on this blog here.

The paper, written in the spirit of Gelman (2022), conducts a thorough re-analysis of my earlier work in light of recent developments in regression discontinuity designs (RDD). It also directly addresses specific critiques of the article raised both in subsequent academic literature and in previous comments on this blog.

A full response to each critique raised on this blog appears in Section A.3 on page 48 of the paper. Among other things, the blog critiqued the use of the global fourth order polynomial, and commenters highlighted that it appeared to be picking up noise in the data rather than a true relationship. While the original article did present results in the SI showing robustness to a local-linear specification with alternative bandwidths, the currentpaper significantly extends these checks and presents new results showing:
  • Polynomial & Bandwidth Robustness: Shows stability across lower-order polynomials, alternative bandwidths, kernel choices, and a donut-hole approach.

  • Noise Reduction via Aggregation: Aggregates data to the level of the running variable to reduce noise, presenting new scatter plots where the visual discontinuity persists at this level of aggregation.

  • Spatial Structure: Demonstrates that reported balance/placebo issues in recent re-analyses stem from ignoring within-constituency clustering and predictive controls.

A link to the paper is here (https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7272598) and the abstract is below:

This note reexamines Thomas (2018), which advances and tests the ‘elite cooperation logic’ whereby national politicians target resources along partisan lines to win over the cooperation of co-partisan state legislators in implementing development projects. Consistent with this logic, a close election regression discontinuity design (RDD) uses project-level data to show that national legislators in North India allocate systematically higher public works expenditures to constituencies of co-partisan state legislators in the period after a state election. Responding to critiques in subsequent literature of the RDD approach used, this note shows that both the evidence of covariate imbalance reported in Bicalho et al. (2026) and the high proportion of significant placebo estimates reported in Albada (2025) are artefacts of ignoring within-constituency clustering and omitting predictive controls. Meanwhile, this note presents re-analyses confirming that the core findings in Thomas (2018) are robust to covariate inclusion using either constituency-level clustering or constituency-level aggregation. Acknowledging the problems related to global fourth order polynomials (Gelman and Imbens, 2019; Albada, 2025), the results in Thomas (2018) are also shown to be robust to lower-order polynomials, alternative bandwidths, alternative kernel choices, the donut hole approach, and an inference procedure that adjusts for worst-case bias (Stommes et al., 2023). The note re-establishes the credibility of the substantive findings in Thomas (2018) and highlights the importance of accounting for spatial clustering in both estimation and diagnostic checks in RDDs.

Anjali took my Bayesian statistics class back in 2006! It’s great to see what former students are doing, and I love seeing this sort of asynchronous collaboration.

Regarding the regression discontinuity issues, I do not think it makes sense to fit unrelated curves on the two sides of the cutoff. So I prefer the versions that fit a single curve plus discontinuity, not two curves.

The other thing I always recommend (see this recent paper with Imbens is that regressions include not just the forcing variable but also other pre-treatment predictors. This should help with both bias and efficiency.

All the discussion of bandwidth, functional forms, etc., can be wasted if the fitted model makes no sense or if it’s a distraction from including additional pre-treatment predictors.

In any case, it’s great to see this sort of open exploration.

Survey Statistics: Modeling Complex Contingency Tables

Andrew looped me into an email thread with folks working on poststratification with partial population information. (He knew I’d be interested, see “poststratification without population level information”.)

They pointed to work by Max Goplerud, Shiro Kuriwaki, Jens Wiederspohn, Adam Conner-Sax, and Philip Greengard, which was recently presented as poster at polmeth: Modeling Complex Contingency Tables.

I wish I got to attend, this is cool ! I’ve got questions about Application 1 (ACS):

  1. What are the covariates w?
  2. Suppose you also have survey data in Palo Alto. Would this be added to the training tables ?
  3. How is the table reconstruction error random ? (I see a density plot.)

The poster gives an answer to Andrew’s question in “Mister P when you don’t have the full poststratification table, you only have margins” and “The continuing challenge of poststratification when we don’t have full joint data on the population”:

I’d recommend first imputing a full poststrat table … But then the question is how to do this.

In Application 2, they ask if a better poststratification table reduces error in estimating the outcome using Multilevel Regression and Poststratification (MRP). Presumably this depends on the distribution of the outcome given the poststratification variables. In “toy example for energy balancing weights” I wrote that raking does well “when Y | X1, X2 is additive”.

I’m excited to read the paper and learn more !

The improvement in political analysis in the past 25 years, as demonstrated by excellent demonstrations of statistical workflow from Elliott Morris, Nate Silver, and Eli Mckown-Dawson

As with baseball, football, and basketball (and I’m sure other sports too), the standard of political analytics is just so much higher than it was, decades ago.

I was talking with Gustavo just the other day about Red State Blue State, and how that work was motivated by confusion following the 2000 election emanating from pundits of the left, right, and center. Back then I felt the compulsion to write a whole damn book to explain what was really going on. I even came up with an entirely new (to the best of my knowledge) concept, “second-order availability bias,” to explain how the journalists could’ve gotten things so wrong.

The concept of “second-order availability bias” never caught on, to say the least: it appears only once in the easily-accessed published literature:

So maybe it’s not such a useful psychological concept. What’s relevant here, though, is that the pundits were getting it so wrong, and with such a consensus, that I felt the need to refute them.

Nowadays, things are different. There aren’t so many all-purpose pundits like David Brooks—people who know essentially nothing and have no real interest in learning but present themselves as infallible experts—, and those who remain don’t have such a platform. Also, with partisan polarization, commentary has become fragmented, and we rarely see much of a consensus among pundits across the political polarization. It’s just not on the table.

Meanwhile, political analytics has become more and more impressive, with important contributions being made by academics, journalists, and political professionals. There’s still disagreement (as here) and some difficulties of communication (as here), but in the past two decades the level has gone up so much: the best analytics has become much more impressive, and what might be called replacement-level analytics has become much better too. I don’t think my own sophistication has increased much at all, and that’s one reason why I’m now less likely to crunch the numbers myself (as I did in the early morning hours of 5 Nov 2008) and more likely to just link to the analyses of others (as with Yair’s report on 2024).

Just today I came across two excellent examples online from journalist colleagues of mine.

Elliott Morris, “I re-analyzed the raw data from Wisconsin’s primary polls. Here’s what actually went wrong,” which features this split-the-difference summary that warms my Bayesian heart:
• Most of the miss in polls in Wisconsin is attributable to faulty demographic targets (too many young people). This inflated Hong’s vote margin by somewhere between 5 and 10 points.
• My best guess is that the race moved 6-10 points toward Crowley after pollsters released their final surveys.
• Non-ignorable non-response within demographic categories likely further inflated Hong’s vote margin by 2-5 points.
Morris goes through lots of details too. I haven’t tried to check any of this, but it seems reasonable. We’ve been saying for a long time that primary elections are hard to predict, but some polls are off by much worse than others, and it’s instructive to look into exactly how this can happen.

Beyond the details and the direct interest of this post to political organizations and pollsters, I appreciate Elliott’s work here because he goes beyond statistical generalities (“regression to the mean,” “sometimes you get a draw from the tail of the distribution,” etc.). This is an important statistical point: the “error term” is only an error term until you drill down, look at more data, and figure out what is going on. It’s so common for researchers to just take their numbers and not think about where they came from (as here)—and, indeed, academics and pundits alike can be rewarded for that sort of asinine don’t-look-carefully-at-the-data attitude. So it’s good to see Elliott demonstrating how it’s possible to do better—if you’re willing to put in the work.

Nate Silver and Eli Mckown-Dawson, “FLIPR 2026 midterm election forecast,” which leads Nate to summarize that “[Michigan Senate candidate] El-Sayed would be an underdog in an election held today and is an underdog in our “Lite” (polls-only) version. The fancy versions look at the fundamentals and are more convinced he’ll come back.”

What I really like about this is how “workflow” it feels. What I’m talking about here is how they fit two different models that are doing two different things, they learn something from the comparison, and then they track this back to their data. This sort of thing isn’t in the textbooks (well, it wasn’t until now) but it’s so important to good applied statistical work. So I love to see it here.

The point about these two posts, one by Morris and one by Silver and Mckown-Dawson, is not that they’re so amazing. I mean, yeah, they’re great, but the real point is how professional they are.

I remember Bill James once wrote, in reaction to the unexpected playoff heroics of Bucky Dent or Ray Knight or somebody like that, that, sure, it’s cool when someone steps up and does the unexpected, but what’s more impressive are those Eddie Murray types who can consistently deliver the expected. Because then you can design a game plan around them and not just have to hope for a miracle.

Morris, Silver, Mckown-Dawson, and others doing what’s now expected, doing it well, and demonstrating modern principles of statistical workflow . . . That’s impressive.

Jamaican me crazy yet again

Regular blog readers will remember Golden Krust, source of our standard unit of currency. (Sorry, Peter Thiel, we don’t take bitcoin.) The above picture is kinda blurry but it gives you a sense of the neighborhood.

But here’s the big news! A new Jamaican beef patty place has appeared, right across the street:

Zooming in, it looks like it’s opening in two days:

I’ll be back soon with a full report (but nothing more on this story).

P.S. We did a taste test! The story is here.