Metascience corner: What can we learn from 12 million empirical research results?

He approves

This is Witold and today we are building a bear.

In our recent paper with Erik van Zwet and Andrew, A statistical case for qualified scientific optimism, we fit metascientific models that use hundreds of thousands of z-values from empirical research. We will be blogging about this paper in the coming months, but the data that we needed for this paper seemed so interesting, that they spun off into a whole separate resource, Benchmarks of Empirical Accuracy in Research. I think readers interested in metascience will find it useful.

BEAR is an open-source compilation of 26 (and counting!) documented metascience datasets. Sometimes it just re-uses data from publications, like Kevin Lang’s work on false positives in economics—and dozen more. (I can’t stress this enough, huge thanks to the legion of researchers who created/maintain these datasets.) For some others I do much more myself, like processing outcomes from tens of thousands trials from clinicaltrials.gov.

And you can access all of that, 12.5 mln results in total, with a few clicks. Just go to https://witold.xyz/BEAR/. BEAR has bespoke datasets from medicine, neuroscience, psychology, ecology, education, economics, political science and more—all standardised and ready to use. Some sets have only z-values, but some have much richer structure: study types, subcategories, effect sizes, type of measures, meta-analytic groupings, etc. Of course quality varies, but I hope the documentation/standardisation will help people pick the right dataset for their research question.

Why does this exist? As I say on the website:

Quantitative metascience drives important debates about research standards. Most of its crucial contributions have been based on analysing individual datasets, often painstakingly constructed by researchers. Thanks to their work we now have better understanding of replicability, publication bias, p-hacking, reporting practices, pre-registration, significance rates, and more. But many canonical findings of metascience are based on individual datasets, each constructed in a specific way. How generalisable are these findings?

In other words, metascience itself can suffer from one of the crucial metascientific concerns: external validity. As someone who works on evidence synthesis I am very concerned with this. To give one example, Erik blogged about the virality of the scary publication bias histogram in biomedical journals. But different metascientific sets give completely different pictures of selection:

I am NOT particularly interested in selection in this post. We could talk about any other metascientific property, e.g. p-hacking, low power. I am also not trying to explain what drives these differences or say that these sets are comparable. (We get into that a little bit in the paper and if people are interested in that particular question, I can write a follow-up post.) The point is that we can learn much, much more when we go beyond one data point. OK, that’s pretty trivial thing to say, but Erik’s example clearly shows that it’s one that hasn’t fully sunk in. So I just hope that BEAR will make it easy to compare metascientific corpora and give people a better picture of the limits of generalisability of various claims. Or generally make metascientific research easier.

For completeness, the GitHub repo also includes the model that Erik, Andrew and I use in the paper I linked earlier. You can use it yourself to e.g. estimate (idealised) expected replication rates, but the model is optional and separate from BEAR data. Big kudos to Erik who got this process started and hunted down many of these sources for the paper. But most importantly the credit goes to teams of researchers who generated and maintain the linked sets.

I also have no doubt that in putting this together I made many mistakes. Even with automated checks, it’s not easy to work with so many data sources in parallel, so I’d be grateful for any corrections!

Do children grow continuously or do they grow in fits and starts?

There’s some literature on saltatory growth in children.

From 1992: Saltation and Stasis: A Model of Human Growth, by Lampl, Veldhuis, and Johnson:

From 1993: A case study of daily growth during adolescence: a single spurt or changes in the dynamics of saltatory growth?, by Lampl and Johnson:

On the other hand:

1993: Linear growth in the rabbit is continuous, not saltatory, by Oerter, Bacher, Cutler, and Baron. This paper unfortunately has no graphs, but the abstract is compelling:

A recent report in Science suggests that human growth occurs in brief bursts, up to 1.65 cm in a single day, separated by extended periods of stasis, lasting up to 63 days. Thus, the organism is proposed to alternate between two states, one with a growth velocity of zero, the other with a mean annualized growth velocity greater than 350 cm/yr. These observations, if correct, suggest the existence of a previously unsuspected hormonal mechanism capable of abruptly switching growth plate cell division on and off and of synchronizing cellular growth not only throughout the growth plate, but presumably throughout all the growth plates in the organism. However, the experimental assessment of short-term growth velocity in the human faces the formidable obstacle of a technical error of measurement that exceeds the mean daily growth rate. Accordingly, we tested the saltatory growth hypothesis by measuring proximal tibial growth in the rabbit, a model in which daily growth rate could be measured more than 15 times more accurately than in the human. The model of saltation and stasis predicts a majority of daily growth velocities clustered around zero, and a minority of high growth velocities, that is, a bimodal distribution. The frequency distribution of observed daily growth velocities instead approximated a single Gaussian distribution, indicating continuous growth. We conclude that linear growth, in the most accurate mammalian system yet studied, is continuous, not saltatory.

A similar analysis was published in 1995, No evidence for saltation in human growth, by Hermanussen and Geiger-Benoit. Again, the paper is not publicly available and I could find no graphs.

The skeptical claims were disputed by a paper from 1996, Is growth saltatory? The usefulness and limitations of frequency distributions in analyzing pulsatile data, which reports:

If the FDGV [frequency distribution of daily growth velocities] is highly skewed, then it is consistent with saltatory growth. However, if the FDGV is not highly skewed, then it is consistent with both the saltatory model and a smooth, continuous growth model, and thus, the results are ambiguous. We conclude that FDGV analysis is not a valid method to exclude saltation and stasis growth processes in longitudinal growth studies.

I came across this published exchange from 1995, where Heinrichs, Munson, Counts, Cutler, and Baron share these data:

to which Lampl, Cameron, Veldhuis, and Johnson reply with this:

There’s also a paper from 2011, Infant head circumference growth is saltatory and coupled to length growth, by Lampl and Johnson, but the graph is not very convincing in that the saltations appear to line up with the precision of measurement:

So here’s my question

Whassup with this saltatory growth thing? It shouldn’t be so hard to study if you just get enough data. It’s not even like you’d need to be measuring a kid for months. Taking several measurements per day for a week or two on a bunch of kids and this should settle it, once and for all, no?

But I couldn’t find much literature on the topic. Do people just believe that the finding was an artifact of bad data collection, so nobody’s followed up on it? The original paper from 1992 has 515 citations on Google scholar, but I couldn’t find any review on the saltation question in particular.

What’s the state of knowledge on this one? It seems weird to have such a specific and measurable question that is still unsettled in this way.

What stories should we tell about science now?

This is Jessica. Like many academics, I am concerned about what sort of new steady state U.S. universities will find themselves in after the dust settles on recent transitions. Namely, the last few years have brought funding cuts, targeted visa policy, reduced demand for grad degrees, and a general brain drain to industry (particularly noticeable in AI and computer science). It’s disorienting to think that academia has already peaked, and that the prestige ranking of the R1 faculty job over the top industry research positions (at least in computer science) might be inverting. But things feel very different than they did even a year ago. The reality of there being less money available to pay for basic aspects of research really started to hit me in the last six months. Post-covid, working on campus became less lively, but now it also feels like our collective attention is anxiously focused on Silicon Valley or Washington D.C. We hold faculty meetings where we discuss things like, Is there any way we can help local faculty members who were laid off from tenure track jobs? How will we ensure we can fund all of our own PhDs, given that TA quotas stay fixed but faculty are running out of funding runway? 

To some, this is an overdue rebalancing. Nate Silver, for example, calls getting a PhD a “much worse value proposition than 20 years ago”, and predicts that elite higher ed will become “~50% less relevant in the new steady state,” which is in his eyes a good recalibration.

But it’s worth reflecting on what is lost exactly, if this dwindling of minds and resources continues. How should we think about the value of what universities provide over industry, like intellectual autonomy, or training on how to think scientifically? As a professor, I could make a list of the things that have kept me in academia–being free to work on the problems I find most important, the diversity of topics I can work on at any given time, grad students who care about doing deep work, having time to think about the best solution to a problem. But at an aggregate level, it’s less clear what the equation is.  

As I was puzzling over all this, I attended a metascience conference, where there was a panel on “the social contract for science.” This is the transactional relationship dating back to at least the 1950s, by which scientists receive public funding and autonomy and society gets the benefits of scientific research. Back in the 19th century, scholars began to make a distinction between “pure” science–research unmotivated by any particular application–and applied science. The social contract takes this distinction and further presupposes a dependence relationship: what is confusingly called the “linear model”, the idea that pure science provides the well from which applied science contributions are drawn. Threaten this foundation, e.g., by letting applied science intercept too much of the resources society puts toward science, and we risk running out of useful innovations. Or so the story goes.

The social contract for science was an attempt to cement the importance of scientific understanding to society, making it an interesting counterpart to the current moment. If the actions of the current administration to direct funds away from universities, and the possibility of using AI to produce research output without understanding, are threatening our sense of what science should be, the history of science policy provides some perspective on how our expectations got shaped in the first place.

In search of the mysterious fruits of basic science

I’ve been reading the work of philosopher Heather Douglas, who has traced and critiqued the basic versus applied science distinction, the linear model as justification, and the idea of scientific freedom as limited social responsibility (see, e.g., here and here, or her book on the value-free ideal). Popularized by Vannevar Bush after WWII, in a report prepared for President Roosevelt, basic science is a reframing of pure science, presented as “scientific capital,” providing the principles and conceptions to power new products and processes years into the future. Bush called for deliberate policy to guard against the otherwise inevitable scenario where applied science drives out the pure. One of the eventual outcomes of his report was the creation of the NSF.

But despite the pragmatic nature of basic science espoused by Bush, as a derivation of pure science, it is hard to separate from less tangible values. One is that scientific understanding is a good outside of practical application, at both the individual and societal level. The earliest advocates of pure science associated it with being closer to God. Post-Enlightenment, this view gave way to a more secular superiority complex, which implied the strong character of the pure scientist, who chose to eschew wealth. “The highest occupation of mankind”, Henry Rowland called it in his Gilded Age era essay, “A Plea for Pure Science,” which bemoaned the vulgarity of attributing scientific greatness to the applied scientist rather than the pure.

From a less moralistic point of view, we’ve been encouraged to believe that a society that has rigorous ways of understanding the world is better off over one that doesn’t. Throughout history, understanding the laws of Nature has been portrayed as a good in itself, along with an intellectual life. From this view, by educating people on how to pursue deep understanding of the world, universities provide the general good of scientific thinking to society. If we believe in the intrinsic value of reading, writing, or intellectual discussion, then it would seem we should value the university as a place that provides the kind of timespan and environment needed to develop these skills.

Some argue that the university has come to serve too many conflicting purposes (research engine, job training center, credentialer, incubator of coming-of-age experiences), and should go back to its classical roots: training in oral reasoning and rhetoric, ethics and moral judgment, historical analysis, and the cultivation of taste and discrimination. This may be a useful refocusing, but it offers little consolation for the fact that the elite research infrastructure that helped this country establish and maintain scientific leadership for decades is in the process of being gutted.

If we take our intuitions from the linear model, we might protest that innovation will suffer if universities’ research purposes are deprioritized. The post WWII science-industrial complex expanded the presence of basic research in industry, but studies suggest that the knowledge generating role of corporate R&D has been on the decline for years. To the extent that basic research is the supplier of downstream applications, it would seem we need universities more than ever.

But the distinction between basic and applied science that’s become synonymous with how we envision science has never been airtight. Critics questioned how an institution could be built around a distinction that seemed to amount to little more than a difference in intention, since applied research sometimes produced important new general knowledge, and pure science contributions sometimes had direct applicability. 

AI research is a recent example. Not only is serious money being made without necessarily requiring advanced degrees, research positions do not require PhDs. By some accounts, passing 30 years old puts one in the older demographic of researchers at frontier AI companies. Yet much of the visible innovation in frontier model development has been heavily concentrated in industry labs, including transformer models, scaling laws, and AlphaFold.

Of course, AI owes much to academia. The amazing thing about deep learning and LLMs, to anyone who was paying attention to NLP before these developments, is that after many years of AI research contributing interesting questions but lackluster results, the technology finally seemed to work. Would we have had the foundations for deep learning if perceptrons had not been stubbornly pursued by academics like Frank Rosenblatt at Cornell early on, picked up again in the 1980s by Rumelhart and McClelland’s Parallel Distributed Processing group, despite multiple periods during which consensus said connectionist approaches were unlikely to pay off?

The challenge is that arguing that “someday the research will pay off,” without being able to point to any hard evidence that basic research is, on average, worth the investment, is not such a convincing argument. According to Douglas, studies have been attempted to show the payoffs of basic research, but without very impressive results. Uncertainty about what time scale we should expect between discovery and application makes this kind of exercise difficult.

At the same time, it’s hard to dispute that monetary incentives can sometimes discourage exploration that would eventually pay off. In evaluating the role of academic research to AI progress, we should keep in mind the uniquely massive private investments AI companies have received, and be cautious using it as a general example. It would be premature to conclude that because progress (in terms of models’ standalone capabilities as measured by benchmarks) doesn’t seem to depend much on academic research at the moment, cutting off academic research would be immaterial. Particularly unfortunate about the historical contingency of frontier companies defining the direction of the field is that they are focused on a pretty narrow space of methods, evaluations, and design ideas. But the power and resources they hold give newcomers to AI research the impression that ideas outside this narrow space aren’t important.

Indulgence, autonomy, and social responsibility

As suggested above, it’s always been tempting to bring moral judgment to bear on the basic versus applied research divide. The latest moment with AI research is no exception. Does the moral high ground belong to those who are staying in academia, underfunded or not, to preserve university culture and their autonomy from corporate interests, or those who are willing to give up a comfortable job to shape the impact of AI as a product in the world?

From one perspective, the academy, as a haven for basic science, has always been at risk of being seen as indulgent. In practice, building a scientific career is in many ways a process of identity development and fulfillment for the scientist. Historical pure science rhetoric associated the pursuit of scientific truth with self-realization. But talking about personal fulfillment does not go over well when your opponent is promoting the idea that science could do more direct good for the country or humanity. Academic scientists have always been at risk of coming off as being self-indulgent, insular, or dilettante when they defend understanding for understanding’s sake. The current political moment is just rehashing old themes.

Another unfortunate historical association of basic science is with insularity and shirking responsibility. After WWI era advances in chemical warfare and explosives, the social responsibility of the scientist became a much greater concern. Philosopher John Dewey came down sharply on the idea that an autonomous space for pure science, unhindered by societal concerns, was something to strive for. Instead, he argued that this impetus to protect pure science was partly a convenient abdication of moral responsibility for the downstream outcomes of research, a “shirking of responsibility.’’

Dewey’s concerns came at a time where philosophy was itself seeking to be more scientific. According to Douglas, Dewey’s views on how philosophers should approach science–through greater integration of societal concerns–lost out to the argument espoused by Bertrand Russell, who instead valorized the “disinterested intellectual curiosity which characterizes the genuine man of science.” The latter view became the more accepted one, and our definition of scientific freedom arose in tandem with expectations of limited social responsibility. It’s not particularly surprising, then, to encounter beliefs that academia is not the place to go if you want to have impact in the world.

At the same time, it seems hard to deny that at this point of time in AI, where we have a large imbalance of power and resources, there is something to be said for the autonomy afforded by the university or nonprofit. Some beliefs about AGI coming out of Silicon Valley border on religious. My biggest concern if I were to join an AI company at this point in time would be losing my ability to think for myself about what problems deserve priority. Having greater agency and impact are attractive, but not if they come at the expense of one’s internal compass or values. As Brendan McCord said recently, “Autonomy is different from agency. Agency is getting things done…You can be more effective than you’ve ever been, and you can be less the author of your own life than you’ve ever been.” Against the groupthink of Silicon Valley, the value of the intellectual autonomy academia provides does feel real. Though it’s unclear how valuable this autonomy will continue to be if academics and others outside the big labs can’t retain enough funding or visibility into frontier model development to remain relevant.

The problem with defining progress as prediction and control 

In a 2014 article called Pure science and the problem of progress, Douglas suggests that if the pure/applied science distinction doesn’t survive scrutiny (which she argues it does not), we’re left with an account of scientific progress based on our ability to predict, intervene, and control our world. But this is not a definition of progress we should be content with:

“Any increase in the capacity to predict or control the thoughts and feelings of human beings would count as scientific progress. An increased capacity to destroy human subpopulations (through, say, targeted pathogens) would count as scientific progress. Developing new heinous capacities would count as scientific progress. Unlike Rowland, we should have no illusions that greater causal efficacy, greater power of intervention, will in fact always provide a better society.” (p. 63)

If misaligned AI, our own creation, changes how we view ourselves and the world, if it convinces some of us it has all of our best interests at heart even as it feeds our insecurities, or pursues its own goals in the background, is that scientific progress?

Douglas argues that judging real progress requires society to weigh in. When it comes to AI, this is happening through pushback against data centers, and the pace of AI progress, and the culture of Silicon Valley. Adoption matters too, but can’t be a substitute for evaluation. We need institutions independent of the companies to help interpret what’s going on. In the midst of changes to so many of our current institutions, we should expect the story we ultimately tell to take time to sort out.

In the meantime, defending academia as a category of research, or a moral standard, is a dead end. What seems more reasonable to advocate is a set of conditions — time, autonomy, training in scientific judgment, the evaluation of new approaches independent of their profitability. The value of these ingredients isn’t easily summarizable in some neat story, because what drives scientific progress is not that simple. But institutions that help society judge what’s been achieved seem worth defending.

One night in Uzbekistan: Why was this one data point so influential, and what should these researchers had done ahead of time to see this?

Sol Hsiang writes:

We have a comment coming out in Nature next week that is going to cause the retraction of a high-profile paper by Kotz et al. from last year (the second most cited climate paper in the news in 2024).

Basically, Kotz et al claimed that climate change was already costing the world economy a huge amount and would cost 300% of what prior estimates claimed (which was already large). This result got enormous attention in Europe, in particular.

We couldn’t reproduce their findings and realized that it was all driven by weird data from Uzbekistan. If you remove Uzbekistan from their data set, the result falls apart. The costs are still large, but not the extreme numbers that made headlines around the world.

One reason we think this is important is because these data were previously being used by central banks around the world to run stress tests for the effects of climate change.

Here’s the retraction note, in full:

The authors have retracted this paper for the following reasons: post-publication, the results were found to be sensitive to the removal of one country, Uzbekistan, where inaccuracies were noted in the underlying economic data for the period 1995–1999. Furthermore, spatial auto-correlation was argued to be relevant for the uncertainty ranges. The authors corrected the data from Uzbekistan for 1995–1999 and controlled for data source transitions and higher-order trends as present in the Uzbekistan data. They also accounted for spatial auto-correlation. These changes led to discrepancies in the estimates for climate damages by mid-century, with an increased uncertainty range (from 11–29% to 6–31%) and a lower probability of damages diverging across emission scenarios by 2050 (from 99% to 90%).

The authors acknowledge that these changes are too substantial for a correction, leading to the retraction of the paper. An updated version of the paper with these changes, which has yet to undergo peer review, is publicly available with continued open access to its data and methodology (https://doi.org/10.5281/zenodo.15984134). The authors intend to submit a revised version of the paper for peer review. If and when published, this retraction note will be updated to include a link to the new publication. The authors appreciate the corrective role of the global scientific community and thank Thomas Bearpark, Dylan Hogan, Solomon Hsiang and Christof Schötz for bringing these issues to their attention. All authors agree to this retraction.

Good for them. And here’s the story in Retraction Watch.

How science advances when data and methods are open

Jonathan Falk independently pointed me to this story and wrote:

Imagine how uphill it would have been without access to the original data/methods.

Good point!

He also pointed to this news article which summarized the story:

If Uzbekistan were excluded . . . the damages would look similar to earlier research. Instead of a 62 percent decline in economic output by 2100 in a world where carbon emissions continue unabated, global output would be reduced by 23 percent. . . .

Wait—the estimate declines by almost a factor of 3 after removing just one data point? Uzbekistan’s not a tiny country but it’s not huge either (population 40 million); it doesn’t seem like its data should have so much influence as all that.

I went back to the original paper and it has some scatterplots, but (a) it’s hard to see that any one point would be so influential, and (b) the countries aren’t labeled so I don’t see which one is Uzbekistan.

A question of influence

What happened with the data? Is there some sort of scatterplot that would’ve indicated a concern?

To put it another way, if the data from a single medium-sized country could have that much of an impact on the findings, that would’ve been worth reporting from the get-go in the original paper. Even had there not been any data problems, we’d want to know that the results were so sensitive to one data point.

So the meta-question is: What data analysis should’ve been done originally, either to flag the problem with Uzbekistan’s data, or at least to reveal the extreme sensitivity of the headline results to that one data point?

I posed this question to Hsiang, who responded:

We noticed this issue because we were looking at several papers and running some basic diagnostics on all of them. One thing we were doing was just dropping one country at a time and rerunning the models to make sure things weren’t being driving by a single country. We were surprised that this turned up. There are many issues with this paper conceptually, but it’s not even really possible to discuss any of them until you deal with the UZB issue. We had a lot of dialogue with the authors, and it turned out that they really hadn’t run much quality control on the more granular data. When we traced back this issue, it seemed like their research assistants had faithfully converted some numbers from a PDF document, but those numbers were just implausible.

There is a scatterplot in their data paper that is supposed to provide technical validation of their data set. We wanted to see why Uzbekistan didn’t jump out, so we reproduced it in our comment (Extended Data Fig 1). It turns out that Uzbekistan wasn’t even the biggest outlier, but that the version they had published had the axes cropped so you couldn’t see the outliers (see red boxes in our version). This seemed indicative of a different issue, which is why we documented it in the comment.

I guess those graphs should be on the log scale?

It still seems crazy that the data from a single mid-sized country could have such a big effect of a global estimate. That’s something that the original researchers should’ve been aware of, and what it suggests to me is that there is a larger methodological problem that this didn’t get looked at automatically during the research process.

Why I am so against volunteering in academia (ecology-edition)

This post is by Lizzie. The photo is from this summer and included for no other reason than because I like a photo with a post. 

A few months ago I wrote a post where I compared the toxic culture of the restaurant industry to the toxic culture of some parts of my academic world in ecology. One point of similarity was how you often need to ‘volunteer’ your time to get a foot in the prestigious door. This led to a query by Phil about why I am against volunteering. Thanks to Phil for asking the query, and others for already effectively giving my general answer, but I will write a full post here since I think it is a good question.

So, why I am so against volunteering in academia? And here I am focusing on my part of academia, which is ecology. In short: I find volunteering in ecology in my world to be an exclusionary practice practiced by those who talk endlessly about inclusion. And, as is often the case in my life, this sort of hypocrisy drives me nuts and I can just never get over it.

And for the longer version….

In my field (ecology) and related fields (evolution, conservation), there is a pervasive assumption that it is a-okay to hire/have ‘volunteers’ to get your work done. I say hire, because these positions are advertised all the time with a list of qualifications you’ll need. Here’s one:

  • BSc degree in a wildlife, environmental or conservation topic or in the process of completing one.
  • Intermediate level in English and Spanish (Oral and Written).
  • Knowledge in wildlife monitoring surveys. Previous research experience in any capacity is a plus.
  • Physically fit and able to work long hours in a difficult and harsh environment.
  • Good team member with excellent communication skills; able to live and work with a multicultural team.
  • Able to live in basic living conditions and tropical rainforest conservation campus.
  • Hard working and passionate with a desire to learn and improve; willing to put the hard work in to go the extra mile for conservation efforts and personal career development.
  • Excellent computer skills with a confidence in all Microsoft and Google programs.

Conservation organizations or conservation-related research positions seem to feel especially allowed to do this, with the argument that their mission somehow absolves them of basic labor laws. The ones that maybe annoy me the most are the ones for ornithology work (that’s a fancy word for the study of birds) where you also need to pay your way to some (often tropical) locale and maybe even your housing and food so you can help slosh around in the jungle and search for birds. If you don’t believe me, most days you can set the Job Type to Volunteer/Training and search for ‘bird’ here and you’ll likely find something like this somewhere:

This is an unpaid position and interns are responsible to cover food and accommodation at the biological station [in Peru/Belize/etc.].

I learned about all this when I was on a ‘field crew’ long ago (that’s a term in ecology for a lot of people working together on some project where they go outside every day to collect data from the natural world). I was doing my PhD on bugs, but everyone else was studying the birds that eat the bugs and they explained to me that usually to get a paying bird-crew job you must first pay your way (housing etc. wherever the field crew is based) to be trained in point counts (standing in one place for a set amount of time and listing all the birds you hear) and then maybe someone pays your room and board and you learn to do ‘nest-finding’ (self-evident definition) and, as you get more and more training, you might get a small allowance (it’s an ‘allowance’ because it is generally way below minimum wage if you do out the per hour rate).

And then maybe you get into grad school to study birds. Lucky you!

And who knows how marine biology works (“Junior scientists in marine mammalogy are expected to have at least one or two unpaid research experiences to qualify for a graduate programme” according to a Nature Careers piece in 2020). Effectively, as best I can tell, the more coveted and sexy the position, the more it depends on extreme levels of unpaid work. I have never seen the data but when I look around at my ornithology and marine biology colleagues I wonder if the t-test on their socioeconomic status before they entered the ornithology/marine biology is higher than other subdisciplines within ecology.

And then the exact same disciples publish articles about how we need more diversity in the sciences. It blows my mind.

For decades people have been pointing out the problem:

Whitaker 2003 The Use of Full-Time Volunteers and Interns by Natural-Resource Professionals (Conservation Biology, Vol. 17, No. 1) explained that most often

full-time [unpaid] positions are filled by aspiring natural-resource professionals (e.g., students or recent graduates) who need workplace experience if they are to advance to graduate school or more lucrative jobs. Consequently, employers often consider work experience adequate compensation for wage shortfalls.

… and continues:

I believe that our widespread use of volunteers and interns to compensate for budget shortfalls does our profession more harm than good. In many cases this approach is in conflict with labor law, hinders the development of new professionals, undermines our profession’s credibility, and is an impediment to achieving our conservation goals.

He then argues that they these positions cause personal hardship, exclude certain groups, and does not meet societal norms:

Our perception of the importance and urgency of our work as conservationists does not elevate us above the societal values that led to [labor laws].’), undervalues conservation and those working in this and related areas.

Our current strategy of cutting wages does not cut costs; rather, it transfers them directly to the lowest tier of professionals in the form of financial and emotional hardship.

Despite this article from 20+ years ago my discipline continues to do this, now at the same time they decry the need for greater diversity in the field. Most articles on how to increase diversity preach for wider acceptance, statements of inclusivity, but not much action on this practice. That said, some articles acknowledge this as a problem we should maybe change if we want to increase diversity. They all cite Fournier & Bond 2015 Volunteer Field Technicians Are Bad for Wildlife Ecology (Wildlife Society Bulletin 39(4):819–821, which is good but I think the fact that they don’t cite Whitaker means they did not try very hard to look into this issue), which to me gives the 3 reasons folks usually give for not paying people who work for them:

  1. I don’t have enough money to pay them.
  2. It’s what has always happened and …
  3. pointing out other worse-treated/compensated workers.

Fournier & Bond 2015 also makes the general argument of why we should not have unpaid interns, with the Biden quote:

Don’t tell me what you value. Show me your budget, and I’ll tell you what you value.

I find (1) pretty common, but (2) is also pervasive. I have tried talking to my colleagues who are very well funded about this and they tell me that this as just ‘how science is done.’ They often explain to me that a student gets a letter of reference from them out of it, so what’s the problem?

I think once it’s engrained that you can get free grunt work, it’s hard to get people to give it up.

And folks are trained in it early. Graduate students are often sent out to get others to help with their grunt work. I see posts on ECOLOG and flyers in my hallway each year explaining who will be needed for a full-time unpaid position to collect soils, or leaf samples or pipette all day long.

What do I think everyone should do? I suggest they start by what I have managed to do: I just don’t allow ‘volunteers’ in my lab and I tell everyone why I don’t allow volunteers. I pay everyone who works in my lab (including undergraduates) and I make sure part of their paid word is training (how to use Git R, sometimes Stan). Is it expensive to pay every undergrad in my lab minimum wage (or better) given my super low budget each year? Hell yes. But that is not an excuse to me — if you cannot afford to get the work done without volunteers, then you cannot afford to get the work done.

Everyone is clearly not going to just do this though, so we could all benefit from some top-down help I suspect. Granting agencies could ask about how much a lab relies on volunteers and discourage it through various mechanisms. Instead, they often encourage it. In Canada, NSERC grades me on how many students I have ‘trained’ (AKA: have passed through my lab) so finding undergraduate volunteers to stock my lab would certainly help the numbers game. Further, NSERC cares about how much you have volunteered in MSc/PhD fellowships, effectively encouraging students to get on board with this practice.

I think NSERC means to support short-term volunteering, which gets back to Phil’s question of ‘when is this okay?’ In part he seems to ask if it is okay to take a minimum wage job on a rec crew or such. To which I say, of course! We’re all welcome to take low-paying jobs if we can get them if you ask me, especially as they are often the gateway to a career change and I support career changes. So the question is more: when is okay to work for free? I work on community-science projects, such as the USA-NPN, where data are entirely collected by volunteers and I see how volunteering in this way or in other ways that help you connect with something beyond your job, give back to community or provide aid are important for a multitude of reasons.

So my focus here is really on full-time or similar positions. Whitaker 2003 makes the distinction between ‘part-time or short-term service’ where you do not have to forego outside opportunities for paid work vs. “other volunteer and internship positions require that individuals live in remote areas and work>=40 hours per week, effectively denying them the opportunity to earn outside wages” and I think this is very good perspective — as it can be extended also cover the undergraduates working in my lab 10 hours/week during term (they do not forego getting part-time wages, which many students need to get through university).

But we should still be thoughtful of the diversity problems volunteering often leads to. Community scientists are often retired and in a good socioeconomic position, and this extends through many other places. If the National Parks Service relies on volunteers to help repair trails etc., I am sure they will have a bias in who signs up for those opportunities and that may not be the best thing for the NPS.

And while I am back on my diversity angle, if you are thinking it is okay for granting agencies like NSERC to ask about volunteering with something along the lines of ‘but don’t worry! If your life is too hard to volunteer, you can just explain that now and we’ll value it,’ then try harder.

Try harder to think through what type of person is the one who will know they just need to profligate themselves and tell all their horrible life problems so the nice middle- and upper-class people on the committee can give them credit for their difficulties. Maybe you get a small slice of people who that works for, but more I think you get several predictable outcomes. You get people lying about it. And you get people who are horrified and scarred by having to do it (IMHO it’s a revolting ask if you back up and think about), or just don’t do it. So while I am on the topic of suggesting we all stop asking people to volunteer to do our work, we should also stop asking people to tell us how hard their lives were.

What does it mean when different articles published by different people on the same topic have nearly identical titles? (“A novel nutritional supplement reduces postprandial glucose response in healthy individuals in a randomised, placebo-controlled, crossover clinical study”)

Sanjeev Sripathi writes:

I came across a firm that seems to be offering a miracle product – https://letsmoderate.com/products/sugar-slayer . Their claim is that ingesting it hammers your post-meal blood sugar spike down by 40%, and they have multiple studies to back it up: https://letsmoderate.com/pages/science

What stood out to me as weird:

(1) These three articles were published at different times by different people but follow the same format for the abstract and text and arrive at nearly the same very large effect size for the same product (the one being sold): https://www.mdpi.com/2072-6643/16/14/2237 , https://link.springer.com/article/10.1186/s41110-024-00275-6 , https://link.springer.com/article/10.1186/s41110-024-00294-3

(2) The other listed articles seem to have been fetched by doing a search for ‘mulberry extract’, which is the critical active ingredient. Whatever I can access does show a notable attenuation on blood glucose and studies exist all the way from 2007 to 2022, thought the impact varies a lot.

(3) These 3 studies are eerily similar to the first 3 I’ve mentioned here: https://www.liebertpub.com/doi/10.1089/jmf.2014.3160, https://link.springer.com/article/10.1186/s12986-021-00571-2, https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0172239#authcontrib . But they’re from 2007, 2015 and 2024, across 3 sets of researchers, which just seems very odd. Is this how everyone is meant to name their papers? i.e. is there a social pressure sitting outside these groups pushing them into common convention?

I’m sure you’d identify more elements if you took a look at it. I’m just unclear if I’m jumping at shadows.

I was curious so I clicked on the link. From the letsmoderate site:

Here are the titles of the first three articles listed in the above email:

A Novel Nutraceutical Supplement Lowers Postprandial Glucose and Insulin Levels upon a Carbohydrate-Rich Meal or Sucrose Drink Intake in Healthy Individuals—A Randomized, Placebo-Controlled, Crossover Feeding Study

A novel nutritional supplement reduces postprandial glucose response in healthy individuals in a randomised, placebo-controlled, crossover clinical study

A ready-to-mix nutraceutical supplement (GLUBLOC) lowers postprandial blood glucose levels in healthy individuals — A randomised, placebo-controlled, crossover study

All the above authors are from India but with no overlap in the author list.

And here are the titles for the other three:

Mulberry Leaf Extract Improves Postprandial Glucose Response in Prediabetic Subjects: A Randomized, Double-Blind Placebo-Controlled Trial

Mulberry leaf extract improves glycaemic response and insulaemic response to sucrose in healthy subjects: results of a randomized, double blind, placebo-controlled study

Mulberry-extract improves glucose tolerance and decreases insulin concentrations in normoglycaemic adults: Results of a randomised double-blind placebo-controlled study

The first one’s from Korea, and the second and third are from England, with some overlap in the author list.

I agree that it’s weird that the papers have nearly identical trials. But I don’t know how things go in the medical literature. Maybe this is standard practice?

I’m not planning to fork over 799 rupees for 30 tablets of Sugar Slayer. One of the article says, “Alkaloid- and polyphenol-rich white mulberry leaf and apple peel extracts have been shown to have potential glucose-lowering effects, benefitting the control of postprandial blood glucose levels,” so maybe I’ll just eat an apple.

Lots of detail on the continuing “Q collar” scandal and why we should probably be doing more coverage of that sort of science-adjacent medical scams and less on repulsively self-promoting but ultimately less harmful academic grifters.

The Q collar, I hope you don’t remember, is a neck accessory that is advertised as “proven” to protect athletes’ brains, but the studies purporting to offer this proof are suspect.

This is an interesting example of science-adjacent incompetence or fraud or grifting or whatever you want to call it, as it overlays junk science with big business (according to journalist Stephanie Lee, the device costs at least $200 and “investment in the Q-collar has totaled more than $30 million”), the government (including the U.S. Army, a U.S. senator, and the Food and Drug Administration), news and social media, and some aspects of the scientific establishment (medical journals).

You can think of Q collar as a baby version of health scams promoted by wealthy and powerful medical-adjacent celebrities such as Dr. Oz, Andrew Huberman, and RFK Jr. It’s junk science, but junk science with a purpose, in this case not so much to promote an ambitious social climber (as with Agus and Arday), it seems to be more about the money. And in that way it could be seen as a more typical (if less dramatic) example of a science scam. It doesn’t involve any plagiarized giraffes or implausible claims of ultra-marathoning, just the exploitation of people’s legitimate health concerns and the lure of the filthy lucre.  It could even be that the scammers believe their own claims (remember, the first step is to fool yourself; the rest then comes easy) and believe they’re doing well be doing good.

We should probably be doing more coverage of science-adjacent medical scams like the “Q collar” and less on the repulsively self-promoting but ultimately less harmful Aguses, Ardays, Wansinks, Hausers, and Arielys of the world who each in his way degrades the reputation of science but isn’t so closely locked in to a moneymaking scheme. (OK. I guess Wansink did ok as a consultant, and Agus and Ariely founded a few businesses which probably netted them a good chunk of change.  They might well be first-class wealthy, even if not in the private-jet stratum.)

Here’s a Q-collar update from science sleuth Mu Yang:

Two new articles on Q collar came out today:

The duplicated tables are almost cute compared to the real issue yet to be addressed. — I mumbled about it in a comment to your post back inJanuary.

Basically, the design is what one would use to teach “confounding factor” in Stats 101:

  • 5 studies (3 on Q collar and 2 on helmet type) reported data from subjects from the same cohort.
  • The subjects in the helmet studies are the controls from the Q collar studies.
  • Two types of helmets were worn by Q-collar-controls, and helmet type had significant impact on brain white matter parameters
  • What helmets Q collar subjects wore is unknown.
  • Q collar is said to affect the same brain white matter parameters as helmet type, in the same direction (the good one)

I [Yang] wrote more about it here and posted on Pubpeer.

Things are at a standstill. The authors have been contacted many times (and I made sure that they are notified about my Pubpeer posts), but remain reticent. It seems that FDA does not care.

Awhile later, Yang added:

Due to the incurable stubborness of some of us, we have been digging deeper in the Q collar junkyard long after Stephanie’s article, your blog post, and the BMJ article.

To put it extremely simply: The control subjects in the Q collar studies were described in two studies on helmet types —which were shown to have substantial impacts on the same brain imaging metrics as Q-collar does. Given controls and Q-collar subjects were students in the same schools, Q collar subjects likely also wore different helmets. BUT, this key confounding factor was never mentioned, quantified, controlled for, or discussed in Q collar papers. I commented on this here. Following that, Dr. M. Shane Tutwiler, a quantitative researcher in the field of educational psychology in the University of Rhode Island, became interested in the case, and wrote a very comprehensive comment on Pubpeer. Shane’s detailed analysis argues that the design and analysis of one of the key studies are grossly inadequate to support the claims of the study—which was crucial for the FDA clearance.

Before posting on PubPeer, Shane spoke to the authors directly to alert them of his concerns and request the data—the 2021 article has a data sharing agreement—but they denied his request because he was critical of their work.

This “silly” and “messy” case is looking much worse than previously thought. However, Dr. Smoliga and I are being stonewalled by SAGE — the six papers under EoC never moved forward. Dr. Tutwiler’s reports to Wiley are also not getting meaningful replies.

We are wondering if you are willing to take a look at the new thread of evidence and weigh in.

My quick answer is that the story seems clear enough already, even without this new thread! I also noticed that one of the authors of one of the papers under discussion is certain “Jeffery N Epstein.” Not the same guy, and I’m guessing that my Columbia colleague Richard “Axel” Foley wouldn’t spend any time responding to his emails.

But I digress.

To return to this horrible “Q collar” story, Shane Tutwiler writes:

I just wanted to close the loop, as it were, on the “Investigation” of the paper I evaluated on PubPeer. The authors shared the data with the editors, and the editors were seemingly content that the paper was perfectly fine as-is (despite many of the issues I pointed out being raised in peer review and ignored by the authors and managing editor at the time). Wiley has closed the case (see below).

If this were a run of the mill journal article that nobody was going to read, I wouldn’t feel so upset. But this research was used by the FDA as part of the evidence chain to justify the device’s use in public. It’s one of those cases where questionable analyses can have real world consequences.

Without any additional knowledge of this particular case, just based on my general understanding of such things, my guess is that the journals’ decision to do nothing here is motivated less by corruption and more by simple laziness. If you retract an article, the authors can scream at you. If you do nothing, people like Tutwiler and me will scream at you . . . but people like us are much less dangerous than authors who refuse to let go of bad research: Those people have money and reputations on the line and can lash out or even sue! So safer to just do nothing and pretend that all is well. The turtle strategy was not so effective recently with Cambridge University, but usually it works just fine, so I guess I can understand why the editors are doing this, even though I think it’s unethical behavior on their part.

And James Smoliga adds some additional context:

The investigation we raised with the Journal of Neurotrauma has been ongoing for over two years now. In the 10 months since the Expressions of Concern were issued, no further action has taken place.

Those six papers are separate from the one that Shane flagged to Wiley… And importantly, that study was the key one that resulted in the FDA’s authorization of the product, which further generated millions in venture capital investment, DoD funding, and consumer purchasing.

This is a mess at every level – improper statistics, inappropriate interpretation, sloppy errors, and allegations of p-hacking (an admission from a collaborator on their research, who doesn’t seem excited to publicly state this).

If you have any interest in writing more about it, or know of any journalists who may be interested in covering this story further, any help is appreciated!

I know lots of journalists, but they usually don’t like covering this sort of story. The outlet most likely to cover it might be Defector, but I don’t know anyone there. Actually, it would fit wonderfully into their Only If You Get Caught podcast. Maybe I can do the contacts-of-contacts thing and see if I can reach anyone there to make the suggestion.

The New York Times already had this article in 2022 that was critical of the Q collar, so I don’t see them running more stories on it. Once you get to the point of junk-science-is-refuted-but-still-keeps-on-making-money, you’re entering “dog bites man” territory, no?

Maybe you could convince noted science skeptics Sean Carroll and Steven Levitt to cover this story . . . ha ha, just kidding! More seriously, maybe If Books Could Kill would be interested, but I don’t know them either.

P.S. More here.

Computer scientists today are in the position of economists in the early 2000s and Freudian psychiatrists in the 1950s

Regular readers of this blog will trace my short career as a Freud expert to a post from 2012, Economics now = Freudian psychology in the 1950s: More on the incoherence of “economics exceptionalism”.

But recently I thought we need to update this analogy.

Back in the early part of this century, economists were riding high: they were the country’s all-purpose pundits, they had tons of influence but were lamenting that they didn’t have enough, and they were going on and on about how special they were, most amusingly in the self-contradictory argument that they were different because they “assume everyone is fundamentally alike; we believe circumstances, not culture, drive people’s decisions.” I’m still not sure what is the difference between “circumstance” and “culture” except that maybe talking about the former is associated with overconfidence.

Nowadays, though, economics is just one more social science. OK, I don’t want to overstate things. I assume they still get paid more than sociologists and political scientists, and, yeah, there’s a Council of Economic Advisers but no Council of Sociology Advisers. Still, I think that economics has lost some of its standing in the past twenty years, partly as a result of the crash of 2008 and its aftermath (political polarization, Brexit, etc.) and partly just the natural ebb and flow of influence, the inevitable cycle of hype and disappointment. Econ hero Steven Levitt was supplanted by data analyst Nate Silver (who identifies as a poker player, not an economist), and we’re not hearing from economists so much anymore, except to hear them fighting in vain against tariffs.

There’s a new hegemonic science in town, and it’s information science, or computer science.

Computer scientists currently have a lot of prestige. Fair enough: they’ve earned it through all the amazing things they’ve built. Indeed, it’s worth comparing to the earlier alpha-dog social scientists. Freudian psychiatry and neoclassical economics were powerful, all-explaining theories that addressed people’s concerns about mental health, happiness, prosperity, and future prospects. They had big theories, which is great. (The theories were not falsifiable, but that can be fine, as the function of such all-encompassing theories is not to make predictions or to explain the world but rather to supply a framework by which the social world can be studied.) Computer science is different: they didn’t develop a theory of society, but they built impressive tools that change how we live in the world.

Now I’d say that computer science today is like economics at the beginning of the century or Freudian psychiatry in the 1950s in being at apex academic and social prestige and influence. Computer scientists, or computer-science-associated businessmen, are the new gurus, in a way that wasn’t the case before. Yes, Steve Jobs and Bill Gates were culture heroes back in the 80s and 90s, but nobody was particularly interested in what they had to say outside of their narrow technological realms. Nowadays, many computer scientists and tech lords present themselves, and are often taken as, all-purpose pundits. (Many of the tech lords are actually tech investors; they’ve absorbed the prestige of information science through the transitive property of money.) As with the economists and Freudians of past eras, they are presented as having some special authority derived in part from their almost inhuman hyper-rationality, a willingness to tell uncomfortable truths.

But then you get the same problem we had before, which is when the gurus and their hangers-on start to believe their own hype.

One of the benefits of living a long time is that you get to see the aftermath of the hype wave. Back when academic economists were on the top of the world, they were complaining that they didn’t have even more power and influence than they already did. Now that economists are just one more group—more influential than political scientists or sociologists to be sure, but no longer apex academics—they’re no longer doing that. It’s when a group is at its most overvalued that it wants even more, and we’re seeing that with the tech industry (and, to extent, their academic allies) today.

Manned Mars Mission Miscellanea

Maciej Cegłowski writes:

Unlike the Moon, which hangs in the sky like a lonely grandparent waiting for someone to visit, Mars leads a rich orbital life of its own and is not always around to entertain the itinerant astronaut. There is just one brief window every 26 months when travel between our two planets is feasible, and this constraint of orbital mechanics is so fundamental that we’ve known since Lindbergh crossed the Atlantic what a mission to Mars must look like. . . .

We shouldn’t send human beings to Mars, at least not anytime soon. Landing on Mars with existing technology would be a destructive, wasteful stunt whose only legacy would be to ruin the greatest natural history experiment in the Solar System. It would no more open a new era of spaceflight than a Phoenician sailor crossing the Atlantic in 500 B.C. would have opened up the New World. And it wouldn’t even be that much fun. . . .

It wasn’t always like this. There was a time when going to Mars made sense, back when astronauts were a cheap and lightweight alternative to costly machinery, and the main concern about finding life on Mars was whether all the trophy pelts could fit in the spacecraft. No one had been in space long enough to discover the degenerative effects of freefall, and it was widely accepted that not just exploration missions, but complicated instruments like space telescopes and weather satellites, were going to need a permanent crew.

But fifty years of progress in miniaturization and software changed the balance between robots and humans in space. Between 1960 and 2020, space probes improved by something like six orders of magnitude, while the technologies of long-duration spaceflight did not. Boiling the water out of urine still looks the same in 2023 as it did in 1960, or for that matter 1060. . . .

Mars is also not the planet we took it for. . . . The surface might be dry, but in most places there was water ice just underneath. Dynamic surface features hinted that water (or at least brine) was flowing to the surface from deep underground. . . . The news from the ground also got better. Arriving at Gale Crater in 2012, the Curiosity rover found itself looking at an ordinary lake bed, complete with organic sediment and odd stick-like structures that would be called fossils if we found them on Earth. The crater had been habitable for millions of years in the past, and something in it was still emitting methane at night. Over in its own crater, the Perseverance rover found complex organic molecules of indeterminate origin.

But the really exciting news for Mars was the discovery of unexpected life on Earth. . . . not just dozens of unsuspected microbial phyla, but two entire new branches of life . . . These new techniques confirmed that earth’s crust is inhabited to a depth of kilometers by a ‘deep biosphere’ of slow-living microbes nourished by geochemical processes and radioactive decay. . . . This underground ecology, which we have barely started to explore, might account for a third of the biomass on earth.

The fact that we failed to notice 99.999% of life on Earth until a few years ago is unsettling and has implications for Mars. The existence of a deep biosphere in particular narrows the habitability gap between our planets to the point where it probably doesn’t exist—there is likely at least one corner of Mars that an Earth organism could call home. . . . if our distant relatives are still alive in some deep Martian cave, then just about the worst way to go looking for them would be to land in a septic spacecraft.

But the fact that a Mars landing stopped making sense has not had the slightest impact on NASA’s plan to go there in a rocket-propelled terrarium.

And more:

The chief technical obstacle to a Mars landing is not propulsion, but a lack of reliable closed-loop life support. . . . The technology program required to close this gap would be remarkably circular, with no benefits outside the field of applied zero gravity zookeeping. The web of Rube Goldberg devices that recycles floating animal waste on the space station has already cost twice its weight in gold and there is little appetite for it here on Earth, where plants do a better job for free. I would compare keeping primates alive in spacecraft to trying to build a jet engine out of raisins. Both are colossal engineering problems, possibly the hardest ever attempted, but it does not follow that they are problems worth solving. In both cases, the difficulty flows from a very specific design constraint, and it’s worth revisiting that constraint one or ten times before starting to perform miracles of engineering. . . . The only way to explore Mars in our lifetime is to ditch the requirement that people accompany the machinery. . . .

In recent years, there’s been a remarkable division in space exploration. On one side of the divide are missions like Curiosity, James Webb, Gaia, or Euclid that are making new discoveries by the day. These projects have clearly defined goals and a formidable record of discovery.

On the other side, there is the International Space Station and the now twenty-year old effort to return Americans to the moon. These projects have no purpose other than perpetuating a human presence in space, and they eat through half the country’s space budget with nothing to show for it. Forget even Mars—we are further from landing on the Moon today than we were in 1965.

In going to Mars, we have a choice about which side of this ledger to be on.

This all makes sense. I’ve never thought much about this Mars mission thing because it’s always seemed like a bit of a joke. But if powerful people are really gonna try to use this as pretext to take a big chunk out of our national resources, then, yeah, it’s good to have people like Cegłowski pushing back.

Why quantitative understanding of effect sizes matters, even if all you care about is the presence of the effect

In reaction to my article with Andy King proposing post-publication review, Dan “Fast and Frugal” Goldstein writes:

Your process limits information search, computation, and time so it seems fast and frugal to me. Happy you still associate me with that term. It was something I coined as a grad student.

In your proposal, only hit papers get audited. It reminds me a bit of Mel Brooks’ The Producers in which the protagonists use the logic “who would audit a flop?” and stay under the radar by intentionally producing a bad show. Fraudsters have likely attempted to make their work seem worthy of publication while ensuring it doesn’t attract too much attention. Just like in The Producers, though, this sometimes backfires.

I don’t know about that! My impression with fraudsters is that they think that fraud is normal science, perhaps out of some mixture of bad education in research methods, a view that “everybody does it,” and a general lack of understanding of how non-cheaters (like you and me!) think. Think of people like Wansink who gave general advice to to the world on how to p-hack, or Gino and Ariely, who published papers on dishonesty, or Mary Rosh, who surely believes that whatever shady statistical manipulations she does are nothing compared to the dastardly deeds done by the Democrats.

I’m sure there’s tons of below-the-radar cheating and bad science that we don’t hear about, but a fair number of prominent science fraudsters seem to enjoy the limelight. One reason for this seemingly self-sabotaging behavior, I think, is that cheating enabled these people to attain great professional success for years. They had no reason to think the juice would stop flowing.

To return to my proposal with Andy King: I think it’s ok that only the hit papers get audited. Bad papers that get no intention aren’t doing much damage, right?

Goldstein adds:

By the way, I was just having a conversation about your sensing that something was amiss with the LaCour study. For years now I have used this quote of yours in a talk I give about putting numbers into perspective. I argue that it’s really important that people learn how to put numbers into perspective because if they don’t, they won’t notice that something is unusual and worthy of a deeper audit. You somehow sensed something was up with the Lacour result. You didn’t think it was fraud yet but you knew it was strange because you know how to put such differences into perspective:

A difference of 0.8 on a five-point scale . . . wow! You rarely see this sort of thing. Just do the math. On a 1-5 scale, the maximum theoretically possible change would be 4. But, considering that lots of people are already at “4” or “5” on the scale, it’s hard to imagine an average change of more than 2. And that would be massive. So we’re talking about a causal effect that’s a full 40% of what is pretty much the maximum change imaginable. Wow, indeed. And, judging by the small standard errors (again, see the graphs above), these effects are real, not obtained by capitalizing on chance or the statistical significance filter or anything like that.

My colleagues and I recently wrote a paper on this general topic of average effect sizes. It’s our contention that people generally are way too optimistic about possible effect sizes, in large part because they don’t think about variation. If you ask someone to hypothesize an effect size, you’ll typically get a guess of the largest effect that might occur.

But what if you don’t really care about effect size–you just want to know about the effect?

For example, maybe you don’t believe that women during certain times of the month are three times more likely to wear red or pink shirts, but you are interested in some sort of evolutionary psychology theory of sexual display. In that case, why should the effect size matter? Why care that a study reported an estimate that was ridiculously implausible?

I have two to this questions, and thus two reasons why effect size is important even for problems where you don’t directly care about effect sizes:

1. Effect sizes vary. An treatment that has an effect (that is, a true effect, not just an estimated effect) of 0.1 for one group of people in one setting could have an effect of -0.2 in some other scenario. A treatment effect in an experiment is the sum of all sorts of things, positive and negative, and there’s no logical reason to think the sign of the effect will be preserved. Effect size matters. The issue is not just that a smaller and more realistic effect size is less important; it’s also that smaller effects can be more easily produced by other factors, and this reduces the generality of any claims, even if the experiment at hand was done well.

2. Experiments produce standard errors as well as estimates. If the standard error from a study is large compared to any realistic effect size, then the study contains very little information. Effect size is important in understanding the informativeness of an experiment, and to do this right you need to have some sense of what the true effect size could be. You can’t just use a point estimate from the study itself, as this estimate will inherently be too noisy to use to judge the information in the study. As I wrote in this note for the Annals of Surgery, Post-hoc power using observed estimate of effect size is too noisy to be useful.

I don’t see journal review as a gatekeeping process that will keep erroneous articles from being published

Someone pointed to this post from last year, “If only Arxiv required researchers to sign at the top rather than the bottom of the page, none of this would’ve happened,” and asked about this statement of mine: “Seriously, though, setting aside the junk references, I don’t know that I would’ve noticed any problems with the paper had it been sent to me cold.”

My corresponded noted that my post pointed out all sorts of statistical analysis problems in the paper, red flags all of the place, and he asked why I would not have right away suspected fraud.

So let me explain.

My remark, “I don’t know that I would’ve noticed any problems with the paper had it been sent to me cold,” reflects that, when I’m sent a paper to review, I don’t review it forensically. Once I see the problems, I can’t un-see them, and, as I wrote in my post, the problems in that paper are clear, but in a quick review I might not have looked into all those details. I don’t think it is a requirement of a reviewer for a journal to investigate a submission in detail. Ultimately the correctness of an article is the responsibility of the author.

To put it another way, I don’t see journal review as a gatekeeping process that will keep erroneous articles from being published; rather, I see it as providing some information to the author and editor.

I’m on record as saying that the problem with peer review is the peers. That doesn’t make peer review useless. It is what it is. Peers can be very helpful in pointing out connections to the literature. I just don’t think it makes sense to see reviewers as gatekeepers. Again, the correctness of the work is ultimately the responsibility of the author. And, in any case, there will be post-publication review of papers that are of interest to later readers.

How generic language shapes the development of social thought

Recently in the sister blog:

Generic language, that is, language that refers to a category as an abstract whole (e.g., ‘Girls like pink’) rather than specific individuals (e.g., ‘This girl likes pink’), is a common means by which children learn about social kinds. Here, we propose that children interpret generics as signaling that their referenced categories are natural, objective, and have distinctive features, and, thus, in the social domain, that such language affects children’s beliefs about the social world in ways that extend far beyond the content they explicitly communicate. On this account, even generics expressing uncontentious content (e.g., ‘Girls are great at math’) can lead children to think of categories as defining fundamentally distinct kinds of people and contribute to the development of stereotypes and other problematic social phenomena.

Here’s the full article.

He “washed his hands in a can of tetraethyl lead at a press conference, claiming he was ‘not taking any chance whatever’. He knew this to be a lie, having already succumbed to a bout of lead poisoning.”

OK, this is absolutely horrifying:

The ill effects of ingested lead and other heavy metals had been known since the 1920s, when employees at TEL [tetraethyl lead]-refining plants began hallucinating butterflies and going into convulsions of violent insanity (at least ten died). ‘Smelter nose’, a finger-sized hole in the septum, was an occupational hazard at plants. Horses near the Bunker Hill stack dropped dead; children were hospitalised with kidney damage, forced to undergo excruciating chelation therapy. By the 1970s scientists were beginning to link lead emissions with surging delinquency and crime rates.

The industry’s response was to deny everything or, at best, occasionally raise the height of its smokestacks. Company quacks put out statements asserting that high levels of lead in human bodies were not only harmless but ‘natural’. Thomas Midgley Jr, a General Motors engineer with the diabolic distinction of having invented both leaded gasoline and chlorofluorocarbons, washed his hands in a can of TEL at a press conference, claiming he was ‘not taking any chance whatever’. He knew this to be a lie, having already succumbed to a bout of lead poisoning. (Years later, paralysed with what was said to be polio, he strangled himself in the ropes of a contraption designed to hoist him out of bed.)

In the 1970s, my dad worked for the EPA in their mobile source enforcement division: their job was to stop people from illegally selling leaded gasoline and to adjudicate petitions from mom-and-pop refineries that, for various reasons, wanted exemptions from the new rules on unleaded gasoline.

But that story about Thomas Midgley, Jr.: Wow. What an evil guy. The linked article (a review by James Lasdun of a book by Caroline Fraser) is just full of horrible stories.

I guess that the world is full of evil people and always will be. The challenge is to avoid putting them in positions where they can do a lot of harm.

Ted-talking University of California professor asks sex trafficker for $3,000,000 because he thinks there’s a “50% chance” he’ll make “important discoveries” in telepathy

This is quite possibly the stupidest thing in the Epstein files.

Lord knows there’s lots of competition from the likes of Soon-Yi “Woody” Allen, Larry “Lawrence” Summers, Nathan “Clippy” Myhrvold, and Columbia’s own Richard “Axel” Foley, but I think this one takes the cake.

In honor of another Epstein associate (see here), I’ll frame it as a “Linda problem”:

Vilayanur is 74 years old, outspoken, and very bright. She majored in neuroscience. As an adult, he was deeply concerned with issues of motor control in stroke victims.

Which is more probable?

1. Vilayanur is a psychology professor who believes in extra sensory perception (ESP)
2. Vilayanur is a psychology professor who believes in ESP and tried to get 3 million bucks from a sex trafficker to fund his lab.

Here’s the story and here’s the evidence:

OK, a dog-bites-man story: over-the-hill professor on the Ted talk cycle seeks return to former glory through the funds of a shady operator.

The interesting part to me is when Ramachandran writes:

In the interest honesty given the focus on such fringe phenomena – theres a 50% chance the whole effort could go up in smoke, but the 50% chance that it COULD lead to something – makes it worthwhile.

I get it that some old guy could think that autistic people have ESP: this kind of mystical thinking was big back in the 1970s when Vilayanur was young, and I think a lot of Boomers have a soft spot for the paranormal. So, sure, why not study the topic–it does no harm.

But to think there’s a “50% chance” that it could lead to something . . . jeez, what an idiot. There’s a big difference between the sane view that doing speculative research on the high-risk, high-reward principle that even if there’s only a 1% chance it comes to something, it could still be worth a shot; and the absolutely insane view that your supernatural study has a 50% chance of being real.

Look, he had tenure, and in any case people have the right to be idiots and not get fired, as long as they do their job well. Ramachandran might well have been an excellent teacher. Or, hey, maybe he was insincere, lying to Epstein in an attempt to scam $3 million from the sex-trafficking financier. But until I hear otherwise, I’ll take Ramachandran’s words at face value and just conclude that he couldn’t think straight about science, I guess in a comparably innumerate way as that physicist who thought that scientific citations were worth $100,000 each. There are a lot of innumerate people out there, and some of them up with tenured academic positions.

It happens. We accept that there are corrupt cops; we also have to accept that there are stupid professors. Some corrupt cops reach heights of political power; similarly, some stupid professors become Ted-talk stars. That’s how it goes.

There’s also the issue, which we’ve discussed before, of which sorts of pseudoscientific beliefs are socially acceptable and which are not. We can laugh or scream at Dr. Oz for promoting astrology (“People that fall under the Aries sign are known for being resourceful, assertive, and headstrong. An Aries can tend to ‘ram or dive in to things head first.’ When an Aries feels blocked, this pent-up energy may appear in the form of migraines, sinus issues, or even jaw tension.”) or out-and-out magic (“You may think magic is make believe but this little bean has scientists saying they’ve found the magic weight loss cure for every body type–it’s green coffee extract.”), but if they believe in the burning bush or the virgin birth or whatever, that’s kind of a different category. Maybe the supernatural stylings of Ramachandran and Oz could be put in the “religion” column and then it would all be cool. I guess the problem is when they try to bring science into the mix rather than just existing on pure belief.

Dr. Oz is currently running the Centers for Medicare & Medicaid Services. That’s a U.S. government agency! I hope he’s not funneling our tax dollars to hucksters promoting “the No. 1 miracle in a bottle to burn your fat,” etc.

More scientists in the Epstein files, including a roboticist and an ESP researcher

I came across this webpage by Sheeva Azma entitled, “Here’s every scientist I have found in the Epstein Files so far.” She’s missing a few big fish:

  • Dan Ariely (professor at MIT and Duke, Ted talk star, and teller of a story about a possibly nonexistent paper shredder)
  • Donald Rubin (professor at Harvard and one of the most influential statisticians of the twentieth century)
  • Stephen Hawking (late physicist and culture hero)
  • Henry Rosovsky (professor and dean at Harvard; ok, he’s just an economist, but some would count this in the “scientist” cattgory)
  • Gerald Edelman (Nobel prizewinning biologist)
  • Stuart Pivar (not an academic but a very successful industrial chemist, so, yes, he counts as a scientist for sure)
  • Jessica Banks (“an inventor, designer, entrepreneur, and roboticist with degrees in Engineering from MIT and Physics from the University of Michigan”)
  • David Gelertner (computer science professor at Yale, most famous as an unfortunate victim of the Unabomber)
  • Roger Schank (cognitive psychologist and all-around asshole who, according to Wikipedia, worked at Stanford University, Yale University, Carnegie Mellon University, and Trump University)

The name that was most interesting to me on Azma’s list was V. S. Ramachandran, listed there as “neuroscientist studying music and the brain.” Many years ago I read a book by Ramachandran, “Phantoms in the Brain,” about his research curing the phantom limb syndrome and related topics. It was really inspirational, but I do remember telling a friend about it at the time and he cautioned me that you can’t always believe what you read in a book, that maybe Ramachandran was exaggerating his successes.

Anyway, Azma links to a news article in the school newspaper of the University of California, San Diego, which reports:

Emails released by the Department of Justice indicate that Jeffrey Epstein provided funding for a UC San Diego lab led by Vilayanur Subramanian Ramachandran, director of UCSD’s department of psychology’s Center for Brain and Cognition and emeritus distinguished professor.

The DOJ released more than 3 million additional pages of the Epstein files on Jan. 30, in which Ramachandran is named by Deepak Chopra, a lifestyle guru with ties to UCSD.

Chopra, a former UCSD family medicine and public health clinical professor, first connected Ramachandran’s lab to Epstein. Chopra told CBS News that he helped Epstein with his struggles with insomnia, including directing him to Ramachandran to learn about ongoing brain research.

OK, fine, nothing wrong so far. A colleague pointed Epstein to this guy’s lab.

But then . . . oh! check this out:

On Sept. 25, 2017, Ramachandran replied to Chopra in an email regarding a study the lab was conducting on an “autistic savant who displays telepathy.” Ramachandran wrote that he does not “have problem with [his] lab being funded by Epstein.”

Ramanchandran further wrote that if Chopra’s “pal [Epstein] is serious about setting in motion a lab for the study of extraordinary brain potential … something like 500,000 to 3 million would get the administrators excited.”

A subsequent email from Epstein to his accountant, Richard Kahn, instructed Kahn to send $25,000 from Epstein’s private foundation, Gratitude America Ltd., to the University of California Board of Regents to fund Ramachandran’s research on savant syndrome. He asked it to be mailed to UCSD’s psychology department’s chief administrative officer, Peter Hinkley, who is still in this position.

Chopra and Epstein’s conversation continues through Oct. 5, 2017, when Chopra updated Epstein on spending the day with Ramachandran to discuss the “pilot study of autistic savants,” confirming their relationship.

This combines several Epstein science themes:
– Junk science (“telepathy”)
– Exploitation of vulnerable people (that “autistic savant” who Ramachandran is using as funding bait)
– Greed (“something like 500,000 to 3 million”)
– Epstein being cagey (he only actually gives $25,000)
– The science-media industrial complex (Deepak Chopra)

The only satisfying thing in all of this is seeing these academic bigshots degrade themselves for so little. According to this website, Ramachandran’s salary was a mere $288,557.00 in the year 2020. I’m actually surprised it’s so low, but, looking it up, I see the source of my confusion. He has medical training and I’d imagined he was in the medical school, but he’s actually just in the psychology department, which doesn’t have the budget to pay med-school-level salaries. But even with his meager under-$300K salary, I’m pretty sure Ramachandran could’ve funded the $25,000 out of his own pocket. But, ohhhhh, that greed . . . he wanted millions!

Is fabricating data worse than fabricating results? Is failing to correct a known false report more or less serious than making the false report in the first place?

Andy King writes:

I have a question for you–and, if you think it worthwhile, for your readers.

A few weeks ago, I was deposed by Harvard’s lawyers in the lawsuit between Francesca Gino and Harvard. Much of the questioning focused on my replications of research by Harvard Business School professor George Serafeim and my allegations of research misconduct against him and his coauthors.

That experience has led to a lively online debate about two questions:
1. Is fabricating data worse than fabricating results?
2. Is failing to correct a known false report more or less serious than making the false report in the first place?

At the moment, my own thinking is this:
1. Both fabricating data and fabricating results mislead readers. They are simply different paths to the same outcome and thus similarly serious.
2. Failing to correct a false report–once the authors know it is false and material–may actually be more serious. It suggests a conscious decision to leave readers with a claim the authors know to be unsupported.

Your ladder of responses to criticism also seems relevant here, especially categories 6 and 7.

Interesting. This has come up in the past, discussing the moral culpability of researchers who make errors and then avoid acknowledging them. For example this guy at the London School of Economics and Political Science, or this guy at the University of Chicago, or, of course, this guy at the University of California. I don’t think that the first two of those people did any direct research misconduct, but they made major research errors that they never acknowledged–they keep pointing to their discredited work without any note of the problems–and, yeah, that seems like misconduct to me.

Here’s another story for ya. Years ago I had a colleague who showed me a paper he’d just written. It read the paper and realized it had a fatal flaw–not a calculation error, but a misapplication or misunderstanding of a statistical model. I won’t go into the details here; what’s relevant to the story right now is that the paper in question had been accepted by the journal but it had not yet been scheduled for publication. This was before the era of online anything, so the paper really was still in process. I told me colleague he was lucky: he could withdraw the paper and spare himself embarrassment. (The error in the analysis was central to the result in the paper; if you got rid of the error, there was nothing to salvage, so it’s not like he could just send in a corrected version.) To my dismay, my colleague replied, No, the paper is accepted, I don’t want to lose a publication. I asked, Doesn’t it bother you to have them publish something that’s wrong?, and he replied something about the literature being self-correcting. I don’t remember the details of this conversation from decades ago, but I do remember the horrible feeling. I thought about contacting the journal to tell them not to publish, but I figured that ultimately it was their problem for accepting it.

A message for Carol Tavris

Dear Dr. Tavris:

I saw in a recent issue of the Times Literary Supplement that you have been critical of the “chambermaid” study which purported to show that people were losing weight without changing their diet or exercise. I agree that this study did not show what it claimed.

Along these lines, you might be interested in two articles I recently published with Nicholas Brown:
How statistical challenges and misreadings of the literature combineto produce unreplicable science: An example from psychology
This is the reason for external replication

Also I looked you up and saw that you were a scholar of feminism, so you might be interested in my post from a few years ago, How feminism has made me a better scientist. Any thoughts on that would be much appreciated.

I was not able to find your email online–for some reason, it’s often hard to find email contacts for people without current university affiliations–so I’m posting this here on the hope that someone who has your contact information can forward it to you.

Yours,

Andrew Gelman
Professor, Department of Statistics
Professor, Department of Political Science
Columbia University, New York

P.S. I blogged the above because I couldn’t find Tavris’s email. But then someone found her email for me. So I emailed her directly. I’ll keep the post up because it could be of interest to others!

P.P.S. The TLS took down Tavris’s review at her request. But it seems to be reprinted here with slight revision.

Turning chaotic sensitivity from a bug into a feature: Using physical modeling and deep learning to alter the paths of storms and mitigate extreme weather events

Qin Huang, Moyan Liu, and Upmanu Lall write:

Extreme weather events, e.g., droughts, floods, heatwaves, and freezes, are increasing in frequency and intensity, posing severe socio-economic impacts as growing populations heighten exposure to risks that conventional infrastructure cannot fully address. We propose supplementing disaster management with Weather Jiu-Jitsu: a strategy that exploits the chaotic sensitivity of mid-latitude atmospheric dynamics to redirect destructive weather trajectories through small, precisely timed perturbations guided by Finite-Time Lyapunov Exponent (FTLE) diagnostics and deep learning forecast models.

They continue:

Proof-of-concept experiments using the Aurora deep-learning Earth system model show that FTLE-guided nudges applied days before peak impact can shift a hurricane track to avoid landfall on a major city, weaken the peak intensity of a blocking-driven cold extreme, and reduce atmospheric river moisture transport under favorable upstream conditions. Control inputs remain below 2% of total system energy in idealized models, though real-world implementation will require advances in monitoring, attribution, and international governance.

There are some cool ideas here. The big ideas are:

1. Small interventions early on can shift the later progression of a storm, and

2. Chaotic unpredictability can be reduced using high-tech machine learning models.

Both these two things are necessary. The first step is needed to allow this to be done with reasonable cost; the second step is needed to give it a good chance of working.

The other cool thing involves cloud seeding. As I understand it, a big hope of the 1950s was idea of seeding clouds to get rain when you want it–but it didn’t really work, because you can’t get it to rain when the water isn’t there. (I’m sure I’m butchering the science here; sorry!) But this new plan is different because you’d be seeding the clouds over the ocean, and the point is not to get it to rain right there but rather to slightly shift where the rain falls.

I can also anticipate political challenges. For example, suppose a storm is headed toward a major city, but if it were diverted it would destroy a resort frequented by rich and powerful people. This is on top of the existing moral hazard by which owners of property near the water expect to be bailed out after natural disasters.

Here are the research papers backing up the idea:

Targeted adaptive chaos control of regimes and eddy strength in two Lorenz models, by Moyan Liu, Qin Huanga, and Upmanu Lall, Chaos, Solitons and Fractals (2026).

Regime identification and control of extremes in the nonautonomous Lorenz model with chaos and intransitivity, by Moyan Liu, Qin Huanga, and Upmanu Lall, Physical Review E (2026).

Upmanu is a water engineer with big ideas. A bunch of years ago he floated the plan to expand Manhattan’s west side by a few hundred meters by taking the silt that is continuously being dredged from the Hudson River and depositing it on the shore as landfill. That never happened but it still seems like a good idea to me. It’s kind of crazy how they’ll spend billions on a single bridge or remodeled train station or whatever but whiff on the big infrastructure projects.

The NIH wants to “Measure and Reward Scientific Impact and Replicable Research Practices.” Here’s my recommendation to the NIH director: you can start by no longer suppressing government reports whose conclusions happen to not be in accord with your ideological preferences.

This came in the email from the U.S. National Institutes of Health:

How Would You Measure and Reward Scientific Impact and Replicable Research Practices?

As NIH continues efforts to strengthen rigor, reproducibility, and public trust in science, we are seeking input from the research community on an important question: Are we measuring and rewarding the activities that matter most for advancing biomedical discovery? NIH wants to hear your perspectives on how scientific impact and rigorous research should be measured and rewarded (NOT-OD-26-087). Comments will be accepted electronically here through our Request for Information (RFI) by August 19, 2026.

My first step would be for the government to stop suppressing its own research. A visible example of this was a report from the Centers for Disease Control and Prevention that appears to have been un-published at the direct orders of the NIH director.

So, yeah, one way to “reward scientific impact and replicable research practices” is to let your own damn employees publish their work.

Beyond that, we have lots of ideas, some of which Erik, Witold, and I discuss in our recent paper, A statistical case for qualified scientific optimism.

P.S. I’m posting this right away, skipping the usual 6-month lag, because the NIH is looking for replies during the next two months.

2015-vintage replication-crisis-era junk science floats into the news

So, I came across this news article titled, “Riley Thinks Suits Make the Coach. Research Says He Might Be Right.”:

The suit had a classic name: the Clark Gable. Navy blue and cut just right, it was the creation of Giorgio Armani, the legendary Italian designer.

It was the piece that made Pat Riley, the legendary NBA coach and executive, believe in the power of style. . . .

“I think an audience wants to see somebody on the sidelines who looks like a leader, dresses like a leader, acts like a leader,” Riley said.

It sounded like a bold claim. Sure, a business suit is undoubtedly nicer than the casual “athleisure” look — team-issue polos and pullovers — that NBA coaches adopted during the COVID-19 pandemic. But can a coat and tie really make someone more of a leader?

“It’s a perfectly reasonable thing to think,” said Abe Rutchick, a professor of psychology at California State University, Northridge. “Which is the idea that the clothes we wear have psychological meaning. We put something on, it’s not just clothes. It means something.”

Uh oh, social psychology research . . .

The article continues:

In the early 2010s, during the rise of casual attire, Rutchick and his colleagues examined a similar question and found something intriguing: Wearing formal attire might actually make a person think and act like a leader.

The researchers, using a variety of cognitive tasks, found that wearing formal clothes caused participants to shift from a concrete mode of thinking to a more abstract mindset — they thought of the big picture and looked further into the future. In other words, they thought like someone who was in charge. . . .

The paper, published in 2015, came a few years after another group of researchers found that people who wore a doctor’s white lab coat — and understood its symbolic meaning — had an increased ability to focus and pay attention. . . .

This sounds pretty bad, no joke. The early 2010s were the high-water mark of junk social psychology. This sort of study was one of the main reasons that the replication crisis became a crisis.

I thought journalists had wised up on this sort of thing, but I guess it remains afloat in the business-inspirational world of leadership.

Don’t get me wrong–I have no problem with these “leadership” stories. It’s cool to read about Pat Riley, and I have no reason to doubt that suit-wearing worked well for him. Everyone has to develop their own personal style. My problem is just with the purported scientific claims.

I found the journal article and, yeah, it’s classic replication crisis fodder:

Study 1: N = 60, p = .03
Study 2: “conceptual replication,” N = 60, p = .05 with 18 people excluded because of missing data
Study 3: N = 34, p = .02
Study 4: N = 54, p = .03 after some data were excluded
Study 5: N = 150, a mix of significant and non-significant results, conclusions made based on whether various inferences reached a significance threshold.

This is pretty much textbook bad statistical analysis of the replication-crisis variety:
– Small sample sizes and noisy data so that there’s essentially no power to detect realistic effect sizes (the kangaroo problem);
– Many researcher degrees of freedom in data exclusion, coding, and analysis, the sort of flexibility that makes it possible to achieve statistically significant p-values even in the absence of any signal;
– A bunch of p-values all in the 0.01 to 0.05 range, which is not what you’d expect from a sampling model of independent experiments (or see here);
– Flexible theories that could explain results through many sorts of interactions (the piranha problem);
– No preregistered replications.

That’s just how they did things back in 2015 so I’m not trying to single out these particular researchers. We know better now. We know not to trust this sort of claims. We don’t need to find a Wansink- or Ariely-style smoking gun; nobody’s suggesting there’s fraud here; it’s just standard-issue junk science of the sort that, until recently, was regularly published in major psychology journals and was regularly featured uncritically in major news media.

The only notable thing to me is to see this sort of claim being pushed in the New York Times now, because I had the vague impression that journalists were now aware of the replication crisis. But I guess there’s still a reservoir of credulity for such claims for stories related to the fuzzy topic of business leadership. I’d hope that straight-up sports reporting would have higher standards for the reporting of research on human performance.

P.S. This is an appropriate post for July 4th now that junk science is ensconced in the U.S. government.