What is the prevalence of bad social science?

Someone pointed me to this post from Jonatan Pallesen:

Frequently, when I [Pallesen] look into a discussed scientific paper, I find out that it is astonishingly bad.

• I looked into Claudine Gay’s 2001 paper to check a specific thing, and I find out that research approach of the paper makes no sense. (https://x.com/jonatanpallesen/status/1740812627163463842)

• I looked into the famous study about how blind auditions increased the number of women in orchestras, and found that the only significant finding is in the opposite direction. (https://x.com/jonatanpallesen/status/1737194396951474216)

• The work of Lisa Cook was being discussed because of her nomination to the fed. @AnechoicMedia_ made a comment pointing out a potential flaw in her most famous study. And indeed, the flaw was immediately obvious and fully disqualifying. (https://x.com/jonatanpallesen/status/1738146566198722922)

• The study showing judges being very affected by hunger? Also useless. (https://x.com/jonatanpallesen/status/1737965798151389225)

These studies do not have minor or subtle flaws. They have flaws that are simple and immediately obvious. I think that anyone, without any expertise in the topics, can read the linked tweets and agree that yes, these are obvious flaws.

I’m not sure what to conclude from this, or what should be done. But it is rather surprising to me to keep finding this.

My quick answer is, at some point you should stop being surprised! Disappointed, maybe, just not surprised.

A key point is that these are not just any papers, they’re papers that have been under discussion for some reason other than their potential problems. Pallesen, or any of us, doesn’t have to go through Psychological Science and PNAS every week looking for the latest outrage. He can just sit in one place, passively consume the news, and encounter a stream of prominent published research papers that have clear and fatal flaws.

Regular readers of this blog will recall dozens more examples of high-profile disasters: the beauty-and-sex-ratio paper, the ESP paper and its even more ridiculous purported replications, the papers on ovulation and clothing and ovulation and voting, himmicanes, air rage, ages ending in 9, the pizzagate oeuvre, the gremlins paper (that was the one that approached the platonic ideal of more corrections than data points), the ridiculously biased estimate of the effects of early-childhood intervention, the air pollution in China paper and all the other regression discontinuity disasters, much of the nudge literature, the voodoo study, the “out of Africa” paper, etc. As we discussed in the context of that last example, all the way back in 2013 (!), the problem is closely related to these papers appearing in top journals:

The authors have an interesting idea and want to explore it. But exploration won’t get you published in the American Economic Review etc. Instead of the explore-and-study paradigm, researchers go with assert-and-defend. They make a very strong claim and keep banging on it, defending their claim with a bunch of analyses to demonstrate its robustness. . . . High-profile social science research aims for proof, not for understanding—and that’s a problem. The incentives favor bold thinking and innovative analysis, and that part is great. But the incentives also favor silly causal claims. . . .

So, to return to the question in the title of this post, how often is this happening? It’s hard for me to say. On one hand, ridiculous claims get more attention; we don’t spend much time talking about boring research of the “Participants reported being hungrier when they walked into the café (mean = 7.38, SD = 2.20) than when they walked out [mean = 1.53, SD = 2.70, F(1, 75) = 107.68, P < 0.001]" variety. On the other hand, do we really think that high-profile papers in top journals are that much worse than the mass of published research?

I expect that some enterprising research team has done some study, taking a random sample of articles published in some journals and then looking at each paper in detail to evaluate its quality. Without that, we can only guess, and I don’t have it in me to hazard a percentage. I’ll just say that it happens a lot—enough so that I don’t think it makes sense to trust social-science studies by default.

My correspondent also pointed me to a recent article in Harvard’s student newspaper, “I Vote on Plagiarism Cases at Harvard College. Gay’s Getting off Easy,” by “An Undergraduate Member of the Harvard College Honor Council,” who writes:

Let’s compare the treatment of Harvard undergraduates suspected of plagiarism with that of their president. . . . A plurality of the Honor Council’s investigations concern plagiarism. . . . when students omit quotation marks and citations, as President Gay did, the sanction is usually one term of probation — a permanent mark on a student’s record. A student on probation is no longer considered in good standing, disqualifying them from opportunities like fellowships and study-abroad programs. Good standing is also required to receive a degree.

What is striking about the allegations of plagiarism against President Gay is that the improprieties are routine and pervasive. She is accused of plagiarism in her dissertation and at least two of her 11 journal articles. . . .

In my experience, when a student is found responsible for multiple separate Honor Code violations, they are generally required to withdraw — i.e., suspended — from the College for two semesters. . . . We have even voted to suspend seniors just about to graduate. . . .

There is one standard for me and my peers and another, much lower standard for our University’s president.

This echoes what Jonathan Bailey has written here and here at his blog Plagiarism Today:

Schools routinely hold their students to a higher and stricter standard when it comes to plagiarism than they handle their faculty and staff. . . .

To give an easy example. In October 2021, W. Franklin Evans, who was then the president of West Liberty University, was caught repeated plagiarizing in speeches he was giving as President. Importantly, it wasn’t past research that was in dispute, it was the work he was doing as president.

However, though the board did vote unanimously to discipline him, they also voted against termination and did not clarify what discipline he was receiving.

He was eventually let go as president, but only after his contract expired two years later. It’s difficult to believe that a student at the school, if faced with a similar pattern of plagiarism in their coursework, would be given that same chance. . . .

The issue also isn’t limited to higher education. In February 2020, Katy Independent School District superintendent Lance Hindt was accused of plagiarism in his dissertation. Though he eventually resigned, the district initially threw their full sport behind Hindt. This included a rally for Lindth that was attended by many of the teachers in the district.

Even after he left, he was given two years of salary and had $25,000 set aside for him if he wanted to file a defamation lawsuit.

There are lots and lots of examples of prominent faculty committing scholarly misconduct and nobody seems to care—or, at least, not enough to do anything about it. In my earlier post on the topic, I mentioned the Harvard and Yale law professors, the USC medical school professor, the Princeton history professor, the George Mason statistics professor, and the Rutgers history professor, none of whom got fired. And I’d completely forgotten about the former president of the American Psychological Association and editor of Perspectives on Psychological Science who misrepresented work he had published and later was forced to retract—but his employer, Cornell University, didn’t seem to care. And the University of California professor who misrepresented data and seems to have suffered no professional consequences. And the Stanford professor who gets hyped by his university while promoting miracle cures and bad studies. And the dean of engineering at the University of Nevada. Not to mention all the university administrators and football coaches who misappropriate funds and then are quietly allowed to leave on golden parachutes.

Another problem is that we rely on the news media to keep these institutions accountable. We have lots of experience with universities (and other organizations) responding to problems by denial; the typical strategy appears to be to lie low and hope the furor will go away, which typically happens in the absence of lots of stories in the major news media. But . . . the news media have their own problems: little problems like NPR consistently hyping junk science and big problems like Fox pushing baseless political conspiracy theories. And if you consider podcasts and Ted talks to be part of “the media,” which I think they are—I guess as part of the entertainment media rather than the news media, but the dividing line is not sharp—then, yeah, a huge chunk of the media is not just susceptible to being fooled by bad science and indulgent of academic misconduct, it actually relies on bad science and academic misconduct to get the wow! stories that bring the clicks.

To return to the main thread of this post: by sanctioning students for scholarly misconduct but letting its faculty and administrators off the hook, Harvard is, unfortunately, following standard practice. The main difference, I guess, is that “president of Harvard” is more prominent than “Princeton history professor” or “Harvard professor of constitutional law” or “president of West Liberty University” or “president of the American Psychological Association” or “UCLA medical school professor” or all the others. The story of the Harvard president stays in the news, while those others all receded from view, allowing the administrators at those institutions to follow the usual plan of minimizing the problem, saying very little, and riding out the storm.

Hey, we just got sidetracked into a discussion of plagiarism. This post was supposed to be about bad research. What can we say about that?

Bad research is different than plagiarism. Obviously, students don’t get kicked out for doing bad research, using wrong statistical methods, losing their data, making claims that defy logic and common sense, claiming to modify a paper shredder that has never existed, etc etc etc. That’s the kind of behavior that, if your final paper also has formatting problems, will get you slammed with a B grade and that’s about it.

When faculty are found to have done bad research, the usual reaction is not to give them a B or to do the administrative equivalent—lowering their salary, perhaps?, or removing them from certain research responsibilities, maybe making them ineligible to apply for grants?—but rather to pretend that nothing happened. The idea is that, once an article has been published, you draw a line under it and move onward. It’s considered in bad taste—Javert-like, even!—to go back and find flaws in papers that are already resting comfortably in someone’s C.V. As Pallesen notes, so often when we do go back and look at those old papers, we find serious flaws. Which brings us to the question in the title of this post.

P.S. The paper by Claudine Gay discussed by Pallesen is here; it was published in 2001. For more on the related technical questions involving the use of ecological regression, I recommend this 2002 article by Michael Herron and Kenneth Shots (link from Pallesen) and my own article with David Park, Steve Ansolabehere, Phil Price, and Lorraine Minnite, “Models, assumptions, and model checking in ecological regressions,” from 2001.

76 thoughts on “What is the prevalence of bad social science?

  1. While I agree with virtually everything Andrew says here, I have misgivings about the Pallesen post. I checked only the second paper he mentioned – the one about blind orchestra auditions. The paper states: “We demonstrate that in the absence of a variable for orchestral “ability,” women fare less well in blind auditions than otherwise. But if the orchestral “ability” of the candidate is held fixed, the screen provides an unambiguous and substantial benefit for women in almost all audition rounds.” The paper is actually quite detailed and has much discussion about lack of “significant” results, as well as a number of “significant” findings. Certainly there are questions to be raised about this particular study. But what I find striking is the simple dismissal of the paper by Pallesen – the paper is anything but simple.

    I didn’t bother to check the other examples because what really bothers me is his use of Twitter (I mean X) comments as evidence. I think those comments almost always oversimplify and are laden with their own methodological problems. I think Pallesen is relying on unreliable, oversimplified, over-hyped, and politicized “evidence.” Andrew’s description of the various methodological dangers and faults is not my issue – I agree with these. But I think oversimplified attacks are not a step in the right direction. The world is complex and statistics is hard. Cherry-picking one table out of a complex paper is not an adequate critique of that work.

    As I said, I did not look at the other papers Pallesen cites, although I am somewhat familiar with Gay’s work and I also think that case is somewhat more complex than Pallesen’s portrayal. I don’t find the plagiarism accusation faulty – but given the political nature of the attacks on Gay, I think that case is more complex than simple differential treatment of administrators/faculty vs students.

    OK, I did look at another of the papers – the Lisa Cook work on patents. It was discussed at length on this blog (https://statmodeling.stat.columbia.edu/2022/02/05/how-many-patents-by-african-americans-were-there-in-the-golden-age-of-innovation-1870-1940/) and I will agree that there was a significant error in measurement that accounted for the most dramatic finding in that paper. But Pallesen’s conclusion that “the flaw was immediately obvious and fully disqualifying” strikes me as oversimplifying and political hype. As the prior discussion on this blog shows, Cook did a lot of painstaking work and raises a number of worthwhile questions – despite the flaw. I’m not sure what that flaw makes “fully disqualifying.” For serving on the Fed? For reading the rest of her paper? All I can say is that this post reminds me about just how bad Twitter/X is for serious discussion.

    • Twitter is pretty crap. Interestingly I’ve found Mastodon to be much less crap. Seems to have been colonized by fairly serious people early on and the culture doesn’t reward the pile-on dunking type analysis, nor does the technology enable it. It feels like there are a lot of people there who are similar to this blog’s commenters.

    • You don’t provide anything to dispute Pallesen’s conclusion. In fact, you actually incorrectly quote him. He does not say that the flaw in Cook’s article is “immediately obvious and fully disqualifying.” What he does say is “It is quite depressing that a paper with such an immediately obvious flaw can achieve such acclaim. ”

      Here is a painstaking reanalysis of her data: https://michaelwiebe.com/assets/cook_reanalysis.pdf

      That “Cook did a lot of painstaking work and raises a number of worthwhile questions” is neither here nor there. especially since she did not do any painstaking work. The lynching data were collected by other people. The data she collected on patents granted blacks from 1870 to 1940 include 726 patents, 400 of which came from a single year, 1900.

      • GB
        Quote directly from above post: “he flaw was immediately obvious and fully disqualifying.” If that is incorrectly quoted, then it was Andrew’s quote of Pallesen. Now that I’ve read some of the critique of Cook’s work, I admit there are serious issues with it. Your statement that “she did not do any painstaking work” does not jive with what I read, however. I will concede that her work is more seriously flawed than I had realized, but I’ll stand by my belief that reducing that work to a total dismissal is not productive. I think most research is not totally good or bad, but somewhere between. Erick Jones may be an exception in that regard, but I don’t think Cook’s work should be so easily dismissed. I feel similarly about the Gay case.

  2. Andrew wrote, “Bad research is different than plagiarism.” That sent me to

    https://www.grammar.com/different-from-vs-different-than

    which is a long discussion regarding when to use “different than” vs. “different from.”

    “So a big distinction between the two expressions is this: different from typically requires a noun or noun form to complete the expression, while different than may be followed by a clause.”

    Historically, “The OED traces the use of different from to 1590 and different than to 1644.”

    And, just to confuse things, “The British also use a structure sounding strange to American ears: different to.” If you do go to the website, be sure to read the comments by the readers.

    • Ah! I didn’t know/realize that “different to” is a Brit thing. I ran into it multiple times over a short period a while ago, and it grated horribly, but all the uses I noticed seemed to be by Americans. I thought my native language had changed out from under me during the 40 years I’ve been hiding under a rock here on the other side of the pond. Thanks!

      It seems “different from”, “different than”, and “different to” are hard to use naturally.

      “Bad research and plagiarism are different.” or “Bad research and plagiarism are different problems.”

      Or “Bad research differs from plagiarism in that it’s not a crime, even though it is.”

      In Japanese, it comes out naturally from the start: The simple “Bad research ha plagiairsm to chigau.” just works. That’s because Japanese is verb final, and you don’t naturally try to do the awkward NP VP NP thing.

  3. In many of the cases Andrew mentions, the problem is the reluctance to admit many areas or questions are not amenable to scientific inference, they are more sciency than science. What’s being measured may be only metaphorically related to what’s claimed to be measured. In some cases, I think it should be admitted that the studies are merely for human interest, or even entertainment value. Rather than concede some fields are fringe, however, there’s a tendency to scapegoat and abandon statistical methods that are valuable in genuine scientific contexts. That’s also bad science.

      • See my response to Dale for clarification. But aside from the fact that I took the topic to concern studies actually carried out, I do think there loads of questions not amenable to being answered by science, e.g., in the realm of metaphysics and ethics.

    • You seem to be implying that a specific type of method (Neyman-Pearson type I’m guessing) is the only true method for doing science, and if that is not being used then it’s not science but only science-like. That seems very convenient for someone promoting a particular type of inference: “well, of course these methods don’t perform well in that situation, because that’s not even science”. That seems rather inadequate for any broader definition of science that covers any type of logical approach to knowing (doesn’t science just come from the Latin “scire”, “to know”?)

      • “You seem to be implying that a specific type of method (Neyman-Pearson type I’m guessing) is the only true method for doing science, and if that is not being used then it’s not science but only science-like.”

        That’s absurd. I never suggested such a thing, nor would I suppose N-P statistics is a “true method for doing science”.

    • I think I can provide an example of what Deborah is referring to – although I’m not saying she agrees with this particular example. I find the whole field of economic experiments an area where what is measured may “only metaphorically” be related to what is claimed. Many experiments are run to see how people behave under various incentive structures, but the generalization to actual behavior requires a leap of faith. To take one notable example, does honesty when signing at the top vs the bottom of a pledge in an experiment tell us anything about honesty in real life situations? I’ve never been convinced that these experiments can tell us anything about people’s real behavior outside of the experimental setup. At the same time, I’ve often used economic experiments in teaching as they are wonderful pedagogical tools to illustrate certain principles. Another more relevant example are survey techniques used to elicit people’s values placed on unmarketed goods, such as good visibility, clean water, etc. I might consider these “not amenable to scientific inference” in that many people have ethical views of these things that don’t (or shouldn’t) get measured by their willingness to pay or accept compensation.

      Now, Deborah may not agree with either of these examples, but they are what her comment makes me think of. The problem I have with her position that the studies are merely for human interest or entertainment value is that these studies attempt to study serious subjects and there may be few alternatives available to study them another way. The proponents of these studies certainly view them like that. If we don’t measure the value of good visibility from such surveys, then how do we compare their value to the values of other things, such as the goods and services forgone in order to provide better visibility? I have an answer to that question – that it is better addressed through civil discourse about ethics and values – but that is a position that is certainly not universally held.

      Perhaps the best example of a question that is “not amenable to scientific inference” is “what questions are not amenable to scientific inference?”

      • Dale:
        I didn’t mean that there’s no way to address some of these questions, sorry. Rather, my point was that many of the actual experiments and ways of analysis renders them sciency. I do not rule out designing an experiment that could be convincing regarding some of these classics. Most importantly, I can think of easy ways to “falsify the study”, as I call it. (Granted, this could wreck some research programs.) Possibly that could put an end to pursuing such experiments in the ways that are now typical (e.g., on cleanliness and morality, political preferences and ovulation–one of Gelman’s examples), or lead to experiments that withstand criticism. If not, then perhaps they should be seen as limited to human interest or entertainment. Those Cosmo articles have a place.

    • Deborah:

      In addition to all of that, I think much of the misunderstanding arises from a lack of quantitative thinking. For example, if someone wants to hypothesize that beautiful parents are more likely to have girl babies, and to come up with an evolutionary story for it, then, sure, why not? Not all science needs to involve experimentation or data; speculation and theory are part of science too. The trouble comes when there’s no sense of the effect size. An effect of 0.001 (which is in the plausible range) is a lot different than an effect of 0.08 (which was observed in the noisy data).

      • Andrew –

        You say:

        > For example, if someone wants to hypothesize that beautiful parents are more likely to have girl babies, and to come up with an evolutionary story for it, then, sure, why not? Not all science needs to involve experimentation or data; speculation and theory are part of science too. The trouble comes when there’s no sense of the effect size

        That’s intersting to me. Is there some way to draw a line between Just-So storification and science? You seem to suggest effect size, but how is that? If so, how is that a criterion that is manifest in real terms?

        Is there no distinction between theory and speculation and “science?” is there a line where they’re a valid part of science and where they’re just useless Just-So stories? Is there no “objective” way to quantify this? I ask this because I see so much “common sense” speculation about evolutionary psychology and the like that looks pretty much like motivated reasoning/confirmation bias and crap to me (see Brett Weinstein).

        • Joshua:

          1) pulling stupid just so ideas out of one’s *** and
          2) calling them “hypotheses” and
          3) testing them with inappropriate methods
          4) hyping the results without any recognition that they’re almost certainly bullshit

          is not science. It’s not scientific speculation. It’s just plain bullshit. The term “hypothesis” is used for tentative explanations of observed phenomena. It’s not a “hypothesis” to ask “are shark attacks in SoCal correlated with the number of daily flushes in the toilets of the Louvre?” That’s just a ridiculous stupid question.

          Second, however, there **is** a mechanism to police junk science:

          1) testing every claim via multiple methods;
          2) stop funding the poeple who continually produce garbage

          What’s funny about social sciences is that people in social sciences seem to think all they have to do is run their little R-package plot up some data and whatever and whatever they conclude is irrefutable scientific law. Sorry, Saturday Morning Cartoon fans, real science takes at least years, usually decades and sometimes centuries, not hours; and normally involves years and years of claims and counter claims as people challenge one another’s claims. It doesn’t just fall out of an hours’ work on a 10×120 data set.

          So, in a sense, right here on this blog you’re experiencing the mechanism that outs junk science: people who don’t agree with it attack it and destroy it.

        • The term “hypothesis” is used for tentative explanations of observed phenomena

          Not really, the term has been diluted. Eg in stats class they will use “hypothesis” to mean “model parameter has the value of exactly zero”.

          The ancient greeks were careful with their terminology. Seems like in the 1800s it started getting sloppy.

          Eg, axiom used to be something so basic people could accept it as obvious. Now it is used to mean *any* assumption you need to make to get the desired result.

          If you look at the “axioms” of ZFC set theory, they almost all look like Euclid’s problematic 5th postulate.

          This matters due to affirming the consequent. Perhaps we can deduce x from either assumptions A, B, C *or* D, E, F. So they are equally good right? Nope, because other consequences may differ. Eg. A, B, C -> y but D, E, F -> z.

          That is why self evident axioms are far better than convenient assumptions. The (accidental?) loss of this distinction seems very problematic to me.

          Similarly, for hypothesis you use the old school definition “an explanation guessed from observation”, this is not happening in social science (and the vast majority of biomed). They don’t even have reliable observations worth explaining.

          The first thing is to go back to something like behaviorism, which was making progress on figuring out how to manipulate conditions so various “laws” of behaviour emerged.

        • Chipmunk –

          I’m not seeing a coherent rule of thumb or a guideline for how to distinguish science from “not science” in what you wrote. It seems that you’re saying it’s a matter of time on task but you don’t provide clear quantitative guideline.

          You speak of social science as if you consider nothing in social science as meriting the label of science.

          So I’m not really getting what you’re saying other than you think some amount of research shouldn’t be funded. How do you think mines should be drawn?

    • “…the problem is the reluctance to admit many areas or questions are not amenable to scientific inference…”

      I fully agree and strongly believe that this problem is an aggressively malignant one in many fields where research can directly impact peoples’ lives. In medicine, the failure to recognize when research methods are incapable of answering an important question *reliably* is the source of massive research money waste and genuine patient harm. Very often, it’s better not to even *attempt* to answer a question if available methods are incapable of providing an answer that’s sufficiently reliable to guide patient care. The length and quality of peoples’ lives can hinge on our willingness to recognize and accept our own ignorance. In short, *some* evidence is very often worse than *no* evidence.

  4. These comments have clarified my thinking a bit. I think what bothers me about the Pallesen quotes, and even Andrew’s characterization, is the good/bad dichotomy. Good/Bad, Left/Right, accept/reject H0, are all binary choices. When it comes to research, it is all too easy to label a research paper as good or bad, or conclude it should never have been done, or label the data as too noisy to yield meaningful results. But I think all of these are a continuum – the world is all gray. When is the data “too noisy?” The real question is “too noisy for what?” I think the problem with these studies is not that the data is too noisy for the study to be done, but that it is too noisy to reach any conclusions. Isn’t the purpose of science to further our understanding, not to reach definite conclusions (although it is nice when we can approximate that)?

    I do find some of these studies to be absurd. I can imagine that the male/female name of hurricanes can be related to their intensity, but I find it absurd to study that. Perhaps that is an example that Mayo might consider for “entertainment” value. Even that absurd question might be worth studying – it can provide a textbook example of spurious correlation. And if an academic undertakes such a study as if it is a serious question, then I think it is appropriate that their annual evaluation/tenure or promotion decisions, take that into account (in a negative way) regardless of whether and where it might get published. But for many of the examples, the noisiness of the data doesn’t render the study worthless – rather it means that no conclusions should be made. As an exploratory analysis, some of these questions (e.g., how effective a particular type of nudge is, to what extent can body language affect performance, etc.) are not worthless, but venturing any conclusions from them is worthless. I don’t see the research as “bad” in itself – what is “bad” is to believe that the research allows you to conclude anything even tentatively. Isn’t all research wrong, but some of it is useful? There is no such thing as a “perfect” study, even if it is an RCT.

    I’m willing to concede that there may be an existence proof of a perfect study from the natural sciences (although I am not sure of this), but I don’t think that affects my point. What I think bothers me about many of these discussions is the binary nature of how we perceive these studies. Pellesen lists studies he seems to believe should never have been done – what I am seeing is mostly studies where the authors have overreached. As a number of people have said, this overreach is encouraged (or even required) by the standards for published work. I am reminded about the Tufte-quoted conclusion (he attributes this to a paper by Bernard Berelson from 1964) from a compilation of the summary findings from 1,045 studies of human behavior:

    “(1) Some do, some don’t.
    (2) The differences aren’t very great.
    (3) It’s more complicated than that.”

    I think that is an adequate description of virtually any study I have seen of human behavior. If these are my expectations, then this should affect my evaluation of a research study. I may find a study absurd that someone else thinks was worth doing – I may believe a laboratory experiment about a nudge with small stakes tells us nothing about how to promote honesty in self-reporting of income, for example, while a researcher may believe the results are somewhat generalizable. It is not that one of us is right and one is wrong. We differ in terms of whether the research is interesting or useful enough to invest energy in, and it is the role of reviewers and editors to decide whether and how the results should be reported.

    So, to return to Andrew’s question in this post: what accounts for the prevalence of bad social science? I would propose that this is the wrong focus – it is the nature of social science that we will differ about what studies are worth undertaking, which are worth reporting, and what conclusions can be reached (or even suggested). I think a better question is what accounts for the failures of our research institutions (including academia, think tanks, granting agencies, etc.) to provide meaningful evaluation of social science research? Until the evaluation improves, I would not expect to see the quality of the work improve. And in a media saturated environment and a world full of misinformation, I’m not optimistic.

    • Imagine a monestary full of monks praying (thinking) and reading trying to invent a smartphone. Further, they only have a vague concept of “smartphone”.

      How many monestaries and how long would it take for them to actually make a smartphone?

      Once you understand what NHST entails, it becomes even more laughable that this method can be used to learn anything than those monks succeeding. It is clear any actual advance must have happened in spite of comparing arbitrary numbers to other arbitrary numbers followed by any combinations of logical fallacies.

      At least the monks will avoid logical fallacies (their issue is with the premises) and think the world should probably make sense.

      • Once again, I find your comment extreme – too extreme. It depends on what you mean by NHST. If all you mean is testing a null hypothesis of no effect, then I agree that the monks praying to invent a smartphone is less dangerous. If you mean all studies that produce p values and attempt to use them in any way, then I think you have gone too far. “It’s more complicated than that.”

        • NHST is the term for testing some default null hypothesis. Similar alternatives to NHST are Fisher’s significance testing and Neyman-Pearson Hypothesis testing.

          In most cases it doesn’t really matter what you use to measure deviation from *your* hypothesis, since they will be about the same and systematic error swamps the “random” error anyway.

        • Anon
          I always find it hard to translate your comments into concrete statements about what you are saying. Here is a concrete case: what is the impact of PSA screening on diagnosis and mortality due to prostate cancer? There are many studies, mostly conducted using standard statistical tools (logistic regression, with and without confounding variables), some RCTs, some observational, mostly with severe selection problems (non-adherence is a big issue in these studies), and with difficult to measure outcomes given the long time frames involved and the frequency of death from other causes even in those with prostate cancer. The list of issues and different studies is long and current thinking is quite uncertain regarding the usefulness of PSA monitoring, biomarker testing, MRIs, and biopsies.

          Are you saying that none of these studies should be done, that all of them should be done, or that only some of them should be done? If the last of these, then which ones and what are you suggesting the criteria are for determining which studies are worth pursuing?

          I think this is an area where systematic errors may well dominate random errors, regardless of whether it is an RCT or an observational study. It is also an area where we have some understanding, though highly imperfect, about the mechanisms that lead to prostate cancer and those that lead to aggressive forms of cancer. Please clarify what you are saying for a concrete application such as this.

        • No one has any idea of whether the PSA tests are a net benefit, or who is benefitting vs getting harmed, and so on.

          The methods used are no more capable of providing this information than the monks are of making a smartphone.

          Every now and then personal experience trumps the “science” like when ICU staff in NYC began refusing to put covid patients on ventilators. But these responses to obvious acute consequences happen in spite of all those kinds of studies.

          When I try to use the literature to make personal health decisions it is almost entirely worthless. There are few things replicated so many times I can take them as roughly a “law” and a few rough principles that make sense but thats it.

          * Thousands of papers have been written about NHST vs significance testing, etc going back to the 1950s. But Gigerenzers “Mindless Statistics” is a good place to start: https://www.sciencedirect.com/science/article/abs/pii/S1053535704000927

          Second would be Meehl’s 1967 “Theory testing in psychology and physics”.

    • Dale:

      I agree there is no sharp line. Indeed, the himmicanes stuff, if real, could yield meaningful practical benefits. If not quite in the “big if true”, it would be in the “moderately big if true” category, at least in the field of disaster preparedness.

      Here’s the thing. If a research team starts a speculative idea, whether it be silly (the idea that people will play better if they’re told they have a “lucky golf ball”) or speculative (the idea that people will react differently to a male or female-named hurricane) or borderline ridiculous (the idea that beautiful parents will be more likely to have girl babies), there are a few ways they can go forward:

      One way to go, which I like, is to advance from a directional hypothesis to a quantitative hypothesis. This takes some work, and I argue it’s work worth doing, as it leads to closer connection to existing science and motivates a critical reading of the literature. It also then leads to the next step of constructing a generative model and simulating fake data that can be used to design a possible experiment.

      Another way to go, which unfortunately seems to still be the standard approach in may areas of psychology research, is to just jump right in and conduct an experiment with no realistic sense of possible effect size and variation, and then find something statistically significant and go publish.

      That second approach will often take everyday mediocre science and turn it into bad science.

      Just for example, here’s an abstract representing mediocre science:

      We speculate that people could react consistently differently to hurricanes with male and female names. This could be studied by comparing death rates in historical hurricanes and further understood using laboratory experiments studying people’s gender-based expectations about severity and preparedness to take protective action.

      And here’s an abstract representing junk science:

      Do people judge hurricane risks in the context of gender-based expectations? We use more than six decades of death rates from US hurricanes to show that feminine-named hurricanes cause significantly more deaths than do masculine-named hurricanes. Laboratory experiments indicate that this is because hurricane names lead to gender-based expectations about severity and this, in turn, guides respondents’ preparedness to take protective action. This finding indicates an unfortunate and unintended consequence of the gendered naming of hurricanes, with important implications for policymakers, media practitioners, and the general public concerning hurricane communication and preparedness.

      The latter is the abstract from the published himmicanes paper; the former is my adaptation of what could’ve been written as speculation. The published abstract is, in my opinion, bad science in combining a lack of strong theory with an absence of evidence. In contrast, my “mediocre” abstract has no strong theory but it does not pretend to have evidence. The addition of the strong and unsupported claims made the project worse.

      Also, the gap between the mediocre and bad abstracts is instructive, in that it suggests the gap, which is some sense of effect sizes and variation, which would help in any study design.

      Of course, the mediocre abstract would never get published in PPNAS or featured in NPR!

      Anyway, to continue to my response to your comment: Papers are of varying quality and there is some continuous range from good to bad; indeed we can see here how a given project can be better or worse, depending on how seriously it is taken.

      • That is helpful. I like the mediocre abstract and the fact that it wouldn’t get published is directly relevant to my question about why our processes for evaluating research are so bad. The idea of looking at “how seriously it is taken” is a nice idea, but rarely practiced. Is it serious because of where it is published or how often it is cited? We know that both of those are highly flawed – I’d suggest the question of how seriously it is taken requires an evaluation – precisely the type of peer evaluation academics try to resist. We do it all the time in classes (especially undergraduate courses), but rarely in peer evaluation. I believe some of such evaluation takes place at the top 20 programs in any field, but that comprises about 1% of the higher education profession. And, even in the top 20 programs, we find too many examples where reputation trumps quality and is extremely resistant to critique.

        To bring in Anon’s response now, prostate cancer screening is taken very seriously judging by how much continued research is done and how much interest there is in the results (like my personal interest). I am as frustrated as Anon is by the difficulty of using any of the published research to inform my personal decisions. But I believe that is due to the nature of the problem, not the fault of the research. I wouldn’t characterize it as no more use than monks making a smartphone, despite the fact that it is difficult to use it for personal or social decisions.

  5. How much of the professor / student double standard simply that professors have more incentive and resources to litigate? I work in a unionized environment, where it is very difficult to discipline a professor even for a gross dereliction; trying to discipline for plagiarism would just be throwing money away. I would guess that even in a non-unionized environment, the potential for litigation plays a role in the decision whether to discipline.

    • Norman,

      Indeed, it can also be difficult to discipline students for cheating. I hadn’t thought about this, but, yeah, that student’s statement, “when a student is found responsible for multiple separate Honor Code violations, they are generally required to withdraw — i.e., suspended — from the College for two semesters,” is probably misleading, for two reasons.

      First, the Harvard president was forced to step down from that position, which as a sanction is perhaps roughly comparable to being suspended as a student for two semesters. So, in that sense, the cases are parallel, and it’s inaccurate to say, “There is one standard for me and my peers and another, much lower standard for our University’s president.” In both cases, there was an investigation and there were consequences.

      Second, selection bias. There are students who, even when confronted with ironclad evidence of cheating, will scream about it and refuse to accept responsibility, and they can convince the honor board or whatever it is to not follow through on any sanctions. I don’t know if they’re threatening to sue or just to raise a fuss, but it can be enough for the university to want to avoid doing anything about it. So, again, I don’t know that the university’s president was being held to a lower standard than students. The main difference seems to be that the discussions regarding the university’s president were held in public, whereas with students, if there’s any public statement it would come at the end of the process.

      • Do you really think that forcing a university president to step down is roughly comparable to suspending a student for two semesters? I don’t (I think the former is a more serious action than the latter). Speculations about how students vs faculty are treated for plagiarism are just that – speculations. It would be interesting if there was some data on this. I know that at some institutions (not Harvard), students are often permitted to be chastised without any formal punishment. It is sometimes viewed as a learning experience where students may not realize exactly when they are plagiarizing. Faculty are thought to understand this and face a higher level of scrutiny – except that faculty are rarely investigated relative to suspicions of student misbehavior. All of this is conjecture on my part, informed by places I have been.

        What I find interesting in many of these recent cases, is that many faculty appear to come from a different culture – their experience seems to not find plagiarism as serious as the critics do, and the standards for authorship and citation seem far more nebulous. I am beginning to wonder if their educational experience has differed greatly from their critics. This is not to condone their behavior – but I think it may be important to understand its causes. If what “we” find so unacceptable has not been part of their experience, then I start to think of it differently. It does not mean I find it ok, but I think taking a hard line based on “our” educational experience may be too simplistic.

        I put “we” and “our” in quotes precisely because I share the educational background of many on this blog. But I am beginning to think that some of the people we have been discussing may have had a very different experience. They possibly view authorship and citation and part of a game whose rules are bendable. While taking a strong stance in favor of unambiguous rules may be appealing – and is what I was trained to believe – it may no be consistent with the experience of other people. How else can I understand some of the egregious, sloppy, or cavalier actions that some of these people engage in? I am inclined to think of these people as bad apples, but I am starting to worry that this is an overly simplistic view. If the system they grew up in is the bad apple, then asserting the superiority of my system over theirs seems an inadequate way to view the situation. I may lament the fact that plagiarism is not valued as highly in some people’s experiences as mine, but I’m not so comfortable with asserting the superiority of my experience (I will admit that I am not so comfortable with the alternative either).

        • Dale,

          In one sense I get where you’re going with this. But what if we applied this same line of thinking to, say, analytical / data handling errors in research that result in misreported results? If a person has been “part of an environment” which views such errors as “part of a game whose rules are bendable” are the errors less significant?

          Some people grow up in a system where crime and theft and corruption are widely condoned. Is it then OK? Because it was tacitly condoned?

          I don’t think so. The fact that some sub-cultures deviate from the widely known and established rules doesn’t mean they don’t know them and should have a get-out-of-jail card. Why then shouldn’t everyone who’s guilty of fakery, intellectual theft or research fraud then just claim “oh, that’s part of the system I was taught” and get off the hook – meanwhile keeping whatever benefits that have come from the inappropriate behavior? And what of the honest people? Why shouldn’t they then turn to using whatever means they need to get what they want? No penalities, right? Or maybe they should be punished for every little slip because they knew better? Yeah, that’s a great barrel of incentives right there. I can see how that would really improve society!

        • Chipmunk:

          Some people grow up in a system where crime and theft and corruption are widely condoned. Is it then OK?

          That seems like a wild extrapolation into a political argument to me! The circumstances around academic misconduct that Dale refers to are rather more subtle.

          For example we recognise (UK) that some students, especially year 1 foreign students with less than optimal English language skills, are at a disadvantage when it comes to writing and may plagiarise with greater or lesser degrees of intent; in some cases they may come from an educational culture where a degree of plagiarism is considered acceptable. We don’t punish our students onerously for this at a first occurrence and consider the consequences, e.g. a discussion with a tutor or Head of teaching, as a learning experience. A serious reoccurence is likely to result in invitation to a plagiarism panel and zero marks for the work but I don’t think anyone considers it appropriate that students should have their degree jeopardized and I would be surprised if US Uni’s are very different (Dale’s comments tend to support this – but not in the case of Harvard apparently!).

          I agree with Dale that cultures (around academic writing in this case) have changed a lot in last 30 years and in fact students can be rather paralyzed by the prospect that they might unknowingly plagiarise in a way that simply didn’t enter our consciousness’s when we were students.

          Very occasionally in a paper written by non-English-speaking scientists I find a sentence or two that I wrote in some paper parroted back to me, and that makes me think that there is probably quite a lot of minor plagiarism around as writers struggle to formulate sentences in their own words in their second or third language. Is this an issue? No I don’t think so and some laxity around very minor degree of plagiarism especially in papers from non-English speakers sems appropriate. Where academic plagiarism is more of an issue is where it is associated with more small-p (or even large-P) “political” subcontexts – e.g. the Wegman example discussed here occasionally. Likewise plagiarism by academics whose plagiarism seems to associate with a journey towards high office is a problem. In these case the professional standards of the plagiarist are rightly called into question.

          IMO minor plagiarism in general academic writing is low on the scale of misdemeanour and might possibly occur without the perpetrator being particulalry aware they”ve made minor coping (the idea that they might accidently do this freaks out some of our students!). On the other hand no-one “inadvertently” white’s-out bands from a western blot and inserts the resulting picture upside-down into a paper figure without knowing that they’re doing something they shouldn’t!

        • chipmunk
          You are a perfect example of what lies at the heart of all the wars about privilege and equal opportunity. In one sense, I agree with you totally – I am not saying that it is ok to violate rules and norms because someone grows up in a different environment. On the other hand, I don’t believe it is sufficient to assert the superiority of “widely known and established rules.” There was a time when the widely known and established rules were that Black people couldn’t eat in the same establishments as Whites; when slavery was widely known and established; when women were widely known and established to do housework and raise children, etc. What I believe is that I am in no place to judge what it means to grow up in an environment very different than my own experience, that I don’t really understand what it means to bend rules if you grow up where the rules seem stacked against you in many ways.

          I don’t reach any conclusions from this. I certainly don’t intend to say behavior that violates these rules is ok and should not be punished. But understanding should imply a more open view than just asserting the superiority of my experience, or yours. I think a lot more needs to be done to explain the reason for these rules and be open to how and why others may not see them the same way.

          Plagiarism is a concrete example. Virtually everyone on this blog – myself included – accepts the importance of rules against plagiarism and the costs associated with violating these. But I am starting to believe that many people don’t see it that way – and I think the right response is to understand better where their beliefs are coming from, rather than simply asserting the superiority of my world view. I am starting to hold similar views for many things – I don’t condone or understand gun violence. But there is a segment of the population that grows up in an environment where that is their reality. I still want laws enforced, but I also think we ultimately need to understand how and why their experience is so different than mine. For one thing, our society is so segregated that life experiences differ markedly from an early age. How does that affect ones’ world views? Do I really understand that? Your position amounts to saying that none of that is relevant or necessary – we need to merely assert our rules and punish those that don’t follow them. Just think about where that eventually leads.

        • Dale, Chris:

          I think it’s safe to say that people don’t throw the book at foriegn students for mild first-offenses of plaigiarism, and I also think it’s safe to say that, whatever the level of corruption in any given society or culture, everyone knows what honesty is. Even capuchin monkeys understand cheating. So, yes, of course there is some gray area around plaigiarism – I don’t reference the Wright Brothers every time I say the word “airplane”. No one expect that.

          As to slavery, it’s hard to understand how anyone in the modern world could be confused by rules that were forcibly ended 150 years ago because they were reviled throughout most of the US and most of the western world. As for segregation, I have never in my life seen it, so whatever might be or have been going on under the table – not very much outside the south during my lifetime – it hasn’t been in the open for half a century.

        • chipmunk
          Now you have confirmed my suspicions that you have lived in a different world than me. If your definition of segregation involves strict legal restrictions on people (e.g., back of the bus, white only golf clubs, etc.), then I agree it has been gone for the past 50 years – though its vestiges are not gone. But when I say segregation, I am referring primarily to neighborhoods in the US. You have to live in a shell to not appreciate how segregated racially most of our lives are in the US. It astounds me if you don’t see that. And, if you do, then it is hard to see that you don’t have any appreciation for the difficulty of understanding how that might shape some people’s realities compared with others.

          As for the understanding of “cheating” I don’t agree that everyone understands its definition. I believe people can “cheat” but feel justified given that they feel they have been “cheated” as well. If you feel victimized by systemic cheating in terms of opportunities, then you may well feel justified cheating on rules set by your perceived perpetrators. I hope we don’t need to debate whether or not such feelings are legitimate – I’m not defending them, I’m trying to understand them. Denying people’s realities does not lead to better understanding.

        • So, yes, of course there is some gray area around plaigiarism – I don’t reference the Wright Brothers every time I say the word “airplane”.

          That’s good to know chipmunk. I guess when you responded to Dale’s comments about some possible sympathetic considerations around motivations for academic plagiarism with “Some people grow up in a system where crime and theft and corruption are widely condoned. Is it then OK?”, you were just extrapolating for effect. Or maybe you like “slippery-slope” arguments. Anyway, yes indeed – there are grey areas around plagiarism.

        • chipmunk
          FYI, I don’t follow Tucker Carlson and don’t even know what that box actually means (other than I know he is right-leaning or on the right fringe, whatever that means). If you want to put me in the “bleeding heart liberal” box, I’ll willingly accept that, though my empathy and policy positions are not the same things.

        • chipmunk –

          > The fact that some sub-cultures deviate from the widely known and established rules doesn’t mean they don’t know them and should have a get-out-of-jail card.

          First, let me ask you whether you think that Americans from a different cultural background than yours are a “subculture?”

          But anyway, surely you’re aware that many students in this country aren’t from the US, and obviously have different cultural backgrounds than the typical American. That means that they may have different communication and educational norms. There are many students for whom the idea of plagiarism from the American contexts seems somewhat illogical. Some students basically see the goal of their time in the classroom as to able to provide the right answer, or basically confirm that they can remember what instruction they were given. As such, it doesn’t particularly matter to them where the ideas come from. The idea of academic work reflecting the individual isn’t part of the cultural norm. The idea of individualism or the idea of “owning” an idea isn’t particularly relevant to them.

          You can have a foreign student who will openly “plagiarize” without knowing that it’s considered “cheating” or dishonest in some way. They can sometimes plagiarize without intending to pass someone else’s work off as their own.

          And of course, there are also a host of other cultural or communicative norms that are considered standard academic norms in American schools but don’t conform to the norms of other cultures.

          Maybe you should consider the contrast between “different cultures” and “subcultures?”

          Since you seem to consider yourself something of an expert on education, perhaps you should look at some of the literature on these topics.

          Here’s a link that might serve as a primer for you (the first hit on a Google search – no vouching for the quality but it lays out some of the general ides).

          https://fixgerald.com/blog/cultural-differences-in-plagiarism#:~:text=Different%20cultures%20have%20different%20customs,and%20how%20cultures%20view%20plagiarism.

        • Joshua
          I don’t know if you intended what sounded to me like a defense of cultural norms that don’t consider plagiarism bad – but that is what is sounded like to me. I see things somewhat differently. I acknowledge that plagiarism is not a shared value in all cultures, but I believe it should be. Where I differ from chipmunk (this is only one of the ways I differ) is that I don’t believe it is productive to assert the superiority of my view over others. Rather, I think the job is that I need to articulate the reasons for avoiding plagiarism and argue for these. In a similar vein, I definitely don’t come from a gun culture, but I realize that many people do. I think my task is then to articulate what I think is bad about that culture and engage with other views.

          To be sure, there are different values in different cultures. Some people lament that fact, some assert the superiority of their view (chipmunk), some consider all values equally valid (you? I’m not sure, but that is the way it struck me), and some (me) see this as an opportunity to engage and understand and perhaps convince.

        • Dale –

          I don’t know if you intended what sounded to me like a defense of cultural norms that don’t consider plagiarism bad – but that is what is sounded like to me. I see things somewhat differently. I acknowledge that plagiarism is not a shared value in all cultures, but I believe it should be.

          To answer, I would need to break down “plagiarism” a bit more. I do think that there’s a moral problem and an educational problem if someone deliberately and deceptively tries to pass someone’s work off as their own. So that’s one type.

          However, I have worked with a lot of brilliant international students – of a type being admitted to universities in the sciences where finding Americans to fill those slots was relatively more difficult – who didn’t think there was anything inherently wrong with presenting work product that wasn’t a output of their individual thought. And that could be considered as another type of plagiarism. (It was part of my job to help those students better understand the American cultural context). As an American educator, I have a strong preference for students to master material and generate work output that is original and “individual,” the idea being that such output is more reflective of a true mastery. But I’m also aware that different paradigms have different strengths and weaknesses and I was working with students who were functioning at levels that enabled them to get coveted positions at prestigious institutions. My point being, that despite by own views about what is a more ideal model, it was undeniable that their model was working quite well by some measures.

          My goal was to help students understand the differences in the models so they could maximize their achievement within an American paradigm, and I will say that as they advanced through their careers they mostly would come to see the advantages of a more individualistic model (as opposed, say, to a more collectivistic one). But even there, my goal was less to communicate that there’s some kind of a hierarchy than to help students choose when and where they would best advance by applying different models.

          To be sure, there are different values in different cultures. Some people lament that fact, some assert the superiority of their view (chipmunk), some consider all values equally valid (you? I’m not sure, but that is the way it struck me),

          It’s not really that I think that all values are equally valid – but in working with a lot of International students and executives it became very clear to me how much “my side” biases can affect how we rank such values.

          As examples, one time I was working with a Japanese executive who was working at ExxonMobil. He was talking about how didn’t see the value in having constant meetings to discuss and produce documents related to the rules and regulations about what could and couldn’t be done in compliance with the company’s ethics. He explained how he saw those efforts as a waste of time as in Japan basically everyone pretty much just knows what the rules are and the point is that there’s a shared and imbedded ethic that was important, rather than that you should be focused on compliance with clearly codified rules. I’m also thinking about one time working with an Indian executive where we discussed differing cultural values about “lying,” and what we discussed is framed by this paragraph I found online about differing cultural attitudes about lying (in this particular paragraph in Korea in comparison to the US):

          America has been characterized as less collectivistic and more individualistic than Korea (Hofstede, 1980, Kim, 1994, Oyserman et al., 2002). Individualists prioritize their own attributes and self-concepts independent of other people, whereas collectivists prioritize interpersonal harmony, fitting in with others, and developing their self-concepts in relation to others (Triandis, 1995). Individualistic cultures emphasize that individuals should have autonomy in their relationships with others (Markus & Kitayama, 1991) and that their personal goals should be respected over group goals (Triandis, McCusker, & Hui, 1990). On the other hand, because individuals in collectivistic cultures consider themselves to be an element of a group, it is more important for these individuals to achieve the group’s goals and ensure their group’s survival rather than to worry about their own personal goals and survival (Markus and Kitayama, 1991, Triandis, 1989). Because of their cultural characteristics, compared to Americans, Koreans may pay more attention to what others may think. Especially when they are to explain their behavioral intentions to others, Koreans may be more likely to use subjective norms as external reasons than Americans may. Thus, it is expected that Koreans and Americans might differ in the extent to which they are likely to use attitudes and norms as external reasons when explaining their behaviors to others.

          https://www.sciencedirect.com/science/article/abs/pii/S0147176711000757

          The point being that where from an individualist lens I might think that lying is objectively always a breach of ethics (as a reflection of me as an individual), someone from another culture might think that telling the truth is an ethical breach if it causes harmful outcomes or disrupts harmony where stretching the truth could have been more beneficial.

          There’s lots more in that article about differing cultural attitudes about “deception.”

          Sorry for wandering kind of far afield…

        • Joshua
          I see some clear issues with both plagiarism and lying (you mention each of these). I’m not so concerned about making sure authors receive credit – it is more about having a clear research record so that we can see how work builds upon past work as well as appreciating the novel aspects. If people come from a cultural background that views that as less important or unimportant, I think it is an important part of the educational process to discuss that – not just so that they understand the “American rules” to follow, but because these are important points. Of course, they have every right to try to convince me otherwise and I’d be interested to hear their reasons.

          Similarly, lying is a problem, especially in this era of mis/dis/no information. Again, if someone comes from a culture where lying is not a problem, there is something to discuss. In fact, I’d say it is critical to the educational process that such discussions take place. I also have considerable experience teaching international students and there are a number of differences from American students. These are always interesting and often the views I was exposed to are not as defensible as the alternatives. But I’d say with lying and plagiarism, I haven’t heard convincing arguments why these are ok, but I am imagining reasons why these might seem less important to some people.

        • Dale –

          It is more about having a clear research record so that we can see how work builds upon past work as well as appreciating the novel aspects. If people come from a cultural background that views that as less important or unimportant, I think it is an important part of the educational process to discuss that – not just so that they understand the “American rules” to follow, but because these are important points. Of course, they have every right to try to convince me otherwise and I’d be interested to hear their reasons.

          I fully agree – again distinguishing between the importance of having a clear record of how the work builds, in contrast to the importance of who it is that did what. So I don’t think it’s a cultural difference with regard to the importance of first aspect but more with regard to the second.

          > But I’d say with lying and plagiarism, I haven’t heard convincing arguments why these are ok, but I am imagining reasons why these might seem less important to some people.

          Again, I’m trying to add more specificity and context to the categories. I don’t know that any culture thinks that lying, per se is OK, but different cultures view context differently. One example would be the subject of a recent popular film related to how its more typical in China to “lie” to people with a terminal illness about the nature of their health. Of course there are cultural differences within the US regarding that issue (in fact, this article about that film makes the claim that attitudes in the US have largely changed over time) but I think it’s probably fair to say that there are broader differences across cultures on some of those types of issues:

          The prevailing narrative of “battling cancer” in Western society has its own issues, with its discourse of personal triumph that values individual responsibility and determination. But the alternative – to lie outright – might seem inconceivable, particularly to those accustomed to the norms of Western culture. It is, however, a common practice in China, rooted in the belief that telling a person about their diagnosis can make their condition deteriorate quicker.

          https://www.newscientist.com/article/2221673-the-farewell-explores-the-ethics-of-lying-about-a-cancer-diagnosis/

      • > First, the Harvard president was forced to step down from that position […] and it’s inaccurate to say, “There is one standard for me and my peers and another, much lower standard for our University’s president.”

        That quote was published on December 31, 2023.

        The position taken by the Harvard Corporation by then had been to reaffirm their unanimous support for her countinued leadership – stating that an independent review of her published work had been conducted by distinguished political scientists finding no violation of Harvard’s standards for research misconduct.

    • A big part of the difference is that common knowledge is different for scholars than for undergraduates. In the latest set of attacks by Rufo he is calling using the full name of a large national study and a straightforward description of its design plagiarism. Also if you are paraphrasing a few sentences into a few sentences there is no need for a citation after every sentence of the paraphrase.

      That said, plenty of students have the resources to litigate, especially at elite schools.

      • Also if you are paraphrasing a few sentences into a few sentences there is no need for a citation after every sentence of the paraphrase.

        I see this sometimes and its really annoying. How does the reader know which preceeding lines are referred to by the ref. And there is an easy solution, just quote the passage. Put it in a footnote if you want.

        Really Id prefer a more modern citation style that points to the exact sentences/figures/datafiles being referenced. I find the majority>/i> of the time the cited source does not actually contain the claimed info. It contains a ref to another source that supposedly does.

  6. > I’ll just say that it happens a lot—enough so that I don’t think it makes sense to trust social-science studies by default.

    I’m glad to see such a frank bayesian prior published by such a prominent person. It should be a wake up call to the entire field.

  7. I expect it would be more awkward to give a citation in a speech than a paper. Thus, we don’t expect them. Which raises the question of why high status people are paid the big bucks to give unvalidated speeches held to even lower standards than papers.

  8. Minor update:
    1) In the 2010 complaints to George Mason University (GMU) versus Ed Wegman&Yasmin Said regarding the Wegman Report (& the overlapping 2007 CSDA paper accepted in a few days by the E-i-C w/o peer review), GMU stonewalled, broke its own rules in multiple ways, gave Wegman a wrist-slap and the Provost even lied to faculty about the process (from FOIAs).
    “See No Evil, Speak Little Truth, Break Rules, Blame Others” including link to 69p PDF with all the details:
    https://www.desmog.com/2012/08/20/see-no-evil-speak-little-truth-break-rules-blame-others/
    Elsivier laudably demanded retraction of the CSDA article in 2011, about 5 months after complaint got to right person.

    2) But in 2011-2012, exposure of plagiarism in two W+S papers published in the WIREs Computational Statististics journal that W+S had started and were Co-Editors-in-Chief led to them disappearing from the masthead without comment. It did take about a year and escalation to the Wiley Board.
    https://www.desmog.com/2015/05/19/ed-wegman-yasmin-said-milt-johns-sue-john-mashey-2-million/

    3) ~2013, GMU administration changed, and complaints to GMU regarding W+S Federal gran misuse seemed to have been taken more seriously, although GMU never replied beyond acknowledging receipt. More plagiarism was documented, but universities are supposed to monitor grants given their %cut, so they may have had to take it more seriously, knowing reports had gone to agencies.
    https://www.desmog.com/2013/05/20/foia-facts-1-more-misdeeds/

    4) In any case, later Wegman was “excluded from the distinguished faculties position which I held” and “was denied to be professor emeritus” when he retired in 2018. (from Wegman deposition in Mann defamation case. He was quite bitter at GMU.)

    SO, GMU should get credit for acting (at least with new President & VP Research/Integrity), but it did take years.

    • Loved seeing that video from Sabine Hossenfelder. It’s a super personal introspection on her career and the reasons why academia failed her and she failed to continue in academia.

    • 15% overhead? Whoa, she is from the 90s! Overhead is 60% or more at most US institutions now init? It requires *alot* of staff with expensive social science degrees to institute DEI programs. Fortunately these are based on the sound science of the social sciences.

      Its great though that this critique comes from someone in the physical sciences. My experience is right on target with hers: I left my PhD (after passing my comps by the way) because I realized that a lot of academia isn’t what it cracks itself up to be. I was startled by how much belief underpins what people think and the ridiculous cred that people give to what’s little more than wild speculation. I thought it was pretty bad in the physical sciences. Then I found the social sciences, which make the physical sciences look like exclusively rock solid truth.

      Dale: Your comment about housing discrimination prompted me to get some background on redlining. It’s tragic. But from my perspective brutally comical that redlining seems to have been instituted and promulgated by the leading Progressives of the 1930s. One can only wonder what horrendous disasters are being schemed up in the government as I write in the name of making things better. The most tragicomic aspect of the story is that it’s hardly the only time government planning has created a disaster. And yet people are completely oblivious to the obvious lesson. (just like most people in academics are oblivous to Hossenfelder’s POV)

      Hossenfelder is on target when she points out why special programs for special groups only institutionalize the discrimination against them. On top of that, they ultimately create a backlash against the groups they claim to protect, since changing how people are selected from a function-relevant performance standard to a function-irrelevant social standard ensures that a high proportion of people unqualified by the functional standard will be selected from the special group based on the social standard. Some of us probably recall how that played out in the 1970s-1980s round of “affirmative action”.

      That’s interesting with respect to social sciences, right? We have a substitution of standard, just like in DEI: people who can get grant money (which entails flogging mythical fantastical discoveries) rather than people who can do sound science. Result? Lots of people who have no idea how to do science are now “social scientists”. At least in name. And the best we can do to try to squeeze some minimal level of utility out of them is create another weak standard that their box-checking minds can grasp: pre-registration.

      • I’ll just also note that the funding scheme in science is analogous to the FHA that created redlinging, right? A centralization of funding authority that kills all the other sources of funding and ensures that whatever biases or failings are within it become institutionalized and very difficult to fix

      • chipmunk
        As usual, your desire/need to make your usual political statements has led you to oversimplify things. Yes, the history of the FHA is sobering – and the misguided policies are worth pondering. But segregated housing was not solely caused by the FHA, nor is its continuation solely due to the FHA and “progressive” forces (https://tcf.org/content/report/economic-fair-housing-act/?agreed=1). The commercial practice of redlining predates the FHA (though the FHA certainly institutionalized it), and zoning laws have been, and continue to be, widely practiced by varied political interests. Some of the “progressives” you so readily dismiss are actively attempting to remove policies that perpetuate segregated housing – while others (progressives as well as whatever you want to call the opposite) continue to engage in exclusionary housing policies.

        I don’t think viewing US housing policy as a left/right issue is sufficient. Nor do I view DEI as a simple matter of “progressive” policy gone awry. I personally hate the institutional facets of DEI – but I acknowledge that there is a real problem that it is a reaction to. You always seem to criticize the misguided (and often destructive) policies while denying that there was a problem to begin with. When pressed, of course, your usual response is to claim that unfettered markets will eliminate any problems.

        Affirmative action and DEI have led to many bad things. But there were bad things before these policies were enacted. Discrimination may contain “the seeds of its own demise” (as expounded by Gary Becker), but many of those seeds have not sprouted. Understanding how and why these policies have failed requires more than blaming those failures on your favorite demonized group.

      • chipmunk –

        changing how people are selected from a function-relevant performance standard to a function-irrelevant social standard…

        You have an amusingly binary pollyannaish/demonizing way of viewing the world.

        What was this world that used an (exclusively) “function-relevant performance standard?” Where did that EVER exist? Do tell.

        It seems that you have a very fixed and cartoonish model for how the world works and then absolutely insist on cramming a messy reality into that model.

    • Dunno.

      I’m not unsympathetic to some of her points about the industry of science, but it’s way too dichotomous and doesn’t describe the reality for many scientists, I am quite sure. Prolly depends on the field to some extent. All feels rather libertarianish, false binaryish (as if there’s anything that doesn’t result in unintended consequences).

      And in the end, the idea that self-promotion on Youtube as an “honest” alternative as opposed to the “dishonesty” and “bullshit” of any scientist who does otherwise seems too circular. As many problems as there are with the current system I have little confidence that we’d be better off if the current system were replaced by a bunch of self-promoters putting out Youtube videos sometimes which I think are fairly hot-takish – for example why is she someone who I’d look to for an evaluation of long covid? Not to say her video on that topic wasn’t interesting or well done, but I don’t think her video is “honest” in contrast to reading someone actively researching and publishing in the field.

      It’s funny to read all the comments on the YouTube video. No doubt there are a lot of scientists who agree with her but hard to think there’s not a selection bias issue.

      This is paralleled by Roger Pielke Jr. on climate science. You proudly stake out a provocative and demeaning argument that will obviously feed the antipathy of groups like contrarians towards scientific institutions, and then play the victim card when the foreseeable reactions (on both sides) occur.

      • I think the “honest” part is in some sense she’s getting paid because people at an individual level want to hear her takes on things, as opposed to getting paid because of some sucking up to institutions and the principal-agent problems they create.

        Sabine is pretty reasonably self aware that she’s not an expert on each and every one of her topics. But people who don’t have PhDs or even Masters degrees or Bachelors degrees in topics like climate science or particle physics or economics or whatever still want someone who has a science background to tell them something *reasonably informed* about those topics. And she does that. Really pretty well I think. I think that’s honest.

        That all being said, I have considered doing YouTube type video series, and in the end I decided not to for a few reasons. One of which is that it’s a tremendous amount of work to write, direct, perform, record, cut, edit, etc the videos together. But I think I could be convinced to do that. What I’m not convinced to do is upload my content to YouTube, a predatory company supported by ads who has the power to make you a rockstar or to cut you off from your audience. Whatever *they* want will happen to you. You are a serf and they are the landlord.

        I’ve considered making a video series and uploading to PeerTube (a decentralized system) but I have yet to decide to do that. I’d probably be willing to do it if I had an income stream to do it, but of course the income streams come with the YouTube brand and little else. Sabine seems like she earns her keep. I hope YouTube deigns to allow her to continue.

      • Daniel –

        I think the “honest” part is in some sense she’s getting paid because people at an individual level want to hear her takes on things, as opposed to getting paid because of some sucking up to institutions and the principal-agent problems they create.

        Sure. At one level that makes sense. At another that’s exactly what someone like John Campbell would say. He puts our some of the worst junk science (covid related) imaginable a couple of times a week on YouTube. He’s got a huge following of people who find his content of value but I’d say he’s getting paid tons of money for complete garbage and outright dishonesty.

        IOW, these are not categorical differences between YouTube science entrepreneurs and academics.

        • Oy. my mother watched him (Campbell) a bunch. I never spent any time looking at him so I hate to hear your assessment. sigh.

          Anyway, no it’s not a categorical YouTube vs Academia, it’s Sabine on YouTube vs Sabine working on Academic grants.

        • Campbell actually started out pretty good. I used to like his videos. Then he took a turn right around the time miraculous claims were being made for ivermectin. He’s in all the way now; died suddenly, turbo cancer, massive blood clots, enormous numbers of excess deaths, vast conspiracies. The whole nine yards.

        • Hmm… You think he’s a true believer, or just found that content to be the content that brings him crazy quantities of cash?

          It’s like Lenin said… you look for the person who will benefit… and … uhh uhh… you know… uhh.

        • There’s a woman, Susan Oliver who was on his podcast, I think a few times, earlier in as an expert. He put out a few things that weren’t correct and she notified him but he never acknowledged. Now she’s put put a whole series of videos focused on his errors and he has never responded. I mean you can never know, but it adds to the likelihood that he’s just grifting.

          He’s now got a huge following and the numbers of subscribers spiked enormously as soon as he started putting out total nonsense. She has highlighted many times how he has said things about the trials or vaccine harms that directly contradict things that he had said early on (things that shouldn’t change over time like like about what he considers an acceptable rate of harm form vaccines).

          He appears to be an earnest believer (it’s part of his whole kind, concerned, grandfatherly persona) and it’s hard to believe someone could intentionally be so blatantly misleading people about their health risks, but really some of it is so crazy and so blatantly inconsistent and his refusal to acknowledge errors pointed out to him… it’s hard to understand how he could not be running a grift.

          I really hate the tendency to attribute malice to a profit motive, because it’s such an unfalsifiable heuristic that can basically be applied to anyone who has a different opinion than your own to explain why they have that different opinion – but sometimes it does truly seem unlikely to be false.

    • That’s an anecdotal example. It does sound like that Dr Hossenfelder had a pretty appalling experience at the institute where she began her independent research with the scholarship that she earned, and the idea that some creep professor would expect his students to write chapters for a textbook he publishes is horrifying.

      But her subsequent scientific experience does rather describe a “first-world problem”. She was able to travel to research abroad (normally a wonderful experience for young scientists although that didn’t work out in the context of her particular personal aims). She discovered that she’s got a talent for preparing science Youtube videos and does this pretty successfully including plugs for commercial organizations (NordVPN; Brilliant).

      And although Daniel says “she failed to continue in academia”, that’s not the case since she’s still affiliated with the University of Munich (Center for Mathematical Philosophy) – if she felt like it she could apply for research grant funding.

      She seems to have the best of both worlds. It’s unfortunate that she didn’t feel able to pursue her dream. I expect that she could have pursued her dream but it might have been difficult for her to do so (science is hard!) and she chose another route. Overall I don’t see her as being a representative example of anything other than Dr Sabine Hossenfelder. Most science graduates don’t continue into academic careers and its useful to have examples of how talented scientists use their scientific knowledge and skills in other spheres.

  9. I thought (and subsequently got confirmation in a comment she made on another science blog) that the specific academic disagreement Dr. Hossenfelder had was mostly with String Theory, not with physics research in general. I have seen many other physicists complain that for the past about thirty years, in order to advance in careers and get published in the field of theoretical particle physics, you had to espouse String Theory.

    I noticed in her blog “Back Reaction” that every few years she moved from one institute affiliation to another, e.g., from Sweden to Canada.

    I found her essays in the that blog interesting and well-written, and I among others suggested she publish them. At first she seemed strongly against taking time away from science to write books, but about a dozen years later (after I had stopped making the suggestion), she got “Lost In Math” published, to some success. Then she got interested in the technical and performance aspects of making videos and started replacing her blog posts with them.

    I subscribe to her videos via Patreon, but have to say I enjoy them much less than the old blog entries which were based on her own work and speculations (and sometimes her twin daughters).

Leave a Reply

Your email address will not be published. Required fields are marked *