Andy King writes:
I have a question for you–and, if you think it worthwhile, for your readers.
A few weeks ago, I was deposed by Harvard’s lawyers in the lawsuit between Francesca Gino and Harvard. Much of the questioning focused on my replications of research by Harvard Business School professor George Serafeim and my allegations of research misconduct against him and his coauthors.
That experience has led to a lively online debate about two questions:
1. Is fabricating data worse than fabricating results?
2. Is failing to correct a known false report more or less serious than making the false report in the first place?At the moment, my own thinking is this:
1. Both fabricating data and fabricating results mislead readers. They are simply different paths to the same outcome and thus similarly serious.
2. Failing to correct a false report–once the authors know it is false and material–may actually be more serious. It suggests a conscious decision to leave readers with a claim the authors know to be unsupported.Your ladder of responses to criticism also seems relevant here, especially categories 6 and 7.
Interesting. This has come up in the past, discussing the moral culpability of researchers who make errors and then avoid acknowledging them. For example this guy at the London School of Economics and Political Science, or this guy at the University of Chicago, or, of course, this guy at the University of California. I don’t think that the first two of those people did any direct research misconduct, but they made major research errors that they never acknowledged–they keep pointing to their discredited work without any note of the problems–and, yeah, that seems like misconduct to me.
Here’s another story for ya. Years ago I had a colleague who showed me a paper he’d just written. It read the paper and realized it had a fatal flaw–not a calculation error, but a misapplication or misunderstanding of a statistical model. I won’t go into the details here; what’s relevant to the story right now is that the paper in question had been accepted by the journal but it had not yet been scheduled for publication. This was before the era of online anything, so the paper really was still in process. I told me colleague he was lucky: he could withdraw the paper and spare himself embarrassment. (The error in the analysis was central to the result in the paper; if you got rid of the error, there was nothing to salvage, so it’s not like he could just send in a corrected version.) To my dismay, my colleague replied, No, the paper is accepted, I don’t want to lose a publication. I asked, Doesn’t it bother you to have them publish something that’s wrong?, and he replied something about the literature being self-correcting. I don’t remember the details of this conversation from decades ago, but I do remember the horrible feeling. I thought about contacting the journal to tell them not to publish, but I figured that ultimately it was their problem for accepting it.
I had a related circumstance many years (also decades) ago. I had completed my PhD and submitted a paper with a clever (I thought so!) bit of chemistry to modify a peptide. The modified peptide had an interesting effect in a cell assay. My supervisor, close to retirement, was happy for me to go ahead and publish by myself.
Around the time the paper was accepted a PhD student still in the lab I’d left contacted me to say they hadn’t been able to reproduce the result I was reporting. I felt like I had no option but to withdraw the manuscript – I wasn’t in a position to go back and redo the work. So I lost a nice single author publication right near the start of my career, though I’m glad I was able to stop it before it appeared in print. Ultimately it doesn’t matter a damn now…
“Is fabricating data worse than fabricating results?“
I think fabricating data is worse. Other people may use your fabricated data without recollecting it, and propagate more errors. Fabricated results can be caught by going back to the original data. Catching fabricated data is harder.
I don’t know that I have a strong feeling, but on its face “It suggests a conscious decision to leave readers with a claim the authors know to be unsupported.” is true for both failing to correct and making the false report in the first place. Unless the false report was an accident, but I don’t think that’s the assumption.
Yeah, I don’t get that either. Fabricating data in the first place is surely also a conscious decision to mislead by definition, arguably a stronger one in that it is a positive decision, whereas failure to acknowledge is negative in the sense that it is the child of inaction, which is more ambiguous morally.
Your moral obligation is to not just to publish results you believe. If that were your only obligation, then fake data and fake results wouldn’t be a problem, so long as you believed them, and you’d never withdraw anything you believed.
Your moral obligation clearly includes true explanations that support your beliefs. Fabricating data or results is a violation of that moral obligation before you even start. Not withdrawing a paper that you know contains false (and uncorrectable) violates the same obligation. If the falsity was through error, you might not feel quite as guilty about the error itself, but your duty is uncjhanged: retract (or cirrect).
This is very complicated stuff in my view, and I reason 1) definitions, and 2) interpretations, and 3) being as clear and specific as possible seem very important. I commented something similar in a recent post. I hope this following example might be useful:
In this case here I don’t know what the exact definitions, and differences between, words like “fabricating data”, “fabricating results”, and “false report” are. This is important for me, for instance, with respect to answering question “2. Is failing to correct a known false report more or less serious than making the false report in the first place?”.
I reason that if “a known false report” is a publication based on fabricated data (e.g. simply come up with data and pretent the data were gathered properly) the scientist is making “a conscious decision to leave readers with a claim the authors know to be unsupported” when publishing in the first place. I assume the “failing to correct a known false report” is not really thought about in this scenario, as it’s unlikely this will be an option the author in the scenario has to ponder. The author already knows it’s a false report from the start, it’s unlikely the author will at one point in time report that it’s a false report, or others will make clear it’s a false report without the author having to “correct” anything in this scenario.
If “a known false report” is a publication based on an error in the data or analyses (e.g. some data were copied incorrectly from the raw data and when this becomes known the subsequent analyses are not valid anymore), I can see the question 2 mentioned above being thought about in this scenario. When the scientist does not attempt to correct this error in this scenario, the scientist is making “a conscious decision to leave readers with a claim the authors know to be unsupported” but in my view this is way less serious than the first example. This is because in the first example of fabricated data:
1) the conscious decision to leave readers with a claim the authors know to be unsupported is also present
2) however the conscious decision to leave readers with a claim the authors know to be unsupported is present from the very start
3) and the conscious decision to leave readers with a claim the authors know to be unsupported is likely a result of a different intention or motive (e.g. intention to deceive in the case of fabricated data from the start, and possibly the intention to not have to deal with embarrassment in the case of an error)
Where does this all stand relative to failing to correct or even doubling down on claims whilst representing oneself as Mr Morality / Open Science?
Here, I am thinking of two episodes amply documented on this blog:
Nelson and Nosek continuing to claim that they set out to examine research practices and replicability long after it was made clear that they set out to examine the decline / supernatural observer effect (only later purging the decline / supernatural observer effect from their published manuscript and published code as preregistered and swapping in the research practices and replicability framing)
Simonsohn supporting and Nelson and Simmons being silent about their P-curve forensic method and its (continued) use by them and others to cast aspersion on literatures and harm careers long after its deleterious statistical properties were made clear.
At least those pointed to in the blogpost do not hold themselves forth as paragons of virtue!
Not:
Regarding your last sentence: I give cheaters zero credit for “not holding themselves forth as paragons of virtue.” What I’d like them to do is the following:
1. Stop cheating.
2. Admit all their previous cheating.
3. Provide restitution for the damage they’ve done.
I don’t give two poops what they hold themselves as.
Quote from above: “Where does this all stand relative to failing to correct or even doubling down on claims whilst representing oneself as Mr Morality / Open Science?”
It could be the case that perhaps one can be influenced too much by what one wants to conclude or show or prove, for instance in designing studies or concluding things from them. At least, I have come to that conclusion after writing my latest manuscript titled “Pre-registration, grocery lists, and particular pre-registration issues” which can be found on SSRN.
When writing that manusript I came across several papers where Mr. Nosek has been involved with. In the manuscript I also refer to the summarizing blog post on this blog titled “What’s the story behind that paper by the Center for Open Science team that just got retracted?” dated september 26th 2024. There is a link to an open science discussion google group post where I post an idea in 2013 and Mr. Nosek reacts, which may provide some additional information about the project if that’s indeed the same project that this retracted paper thing concerns.
I always try and keep the option open that one may never know what the reasons are for certain things, or who is responsible for what. But, as a scientist an author, I also think it’s okay to work with what can be seen and read, which is what I have attempted to do concerning my manuscript. I end my manuscript with the following, which I reason also applies to readers of my manuscript:
“It might be best to let readers direct their attention to that package of work and results that provides them with the sort of information and perspectives they need (see Heesen & Bright, 2021, pp. 649-650), to let readers check and verify themselves, and to let readers draw their own conclusions (see Scott & Jones, 2017, p. 2219).”
Quote from above: “Where does this all stand relative to failing to correct or even doubling down on claims whilst representing oneself as Mr Morality / Open Science?”
I have started to ponder and wonder more and more about the following option. What if one promotes certain things, or emphasizes certain things, largely as a way to try and accomplish something related but different down the road.
For instance, organizing a big replication project where you ask many other researchers to collaboratively replicate studies may also promote a certain website or platform if that website or platform is used by these other researchers to replicate the studies. In this way, the big collaborative replication project is also kind of a big advertisement for the website or platform, especially when newspapers and tv stations provide additional attention and let the owner of the website or platform speak about certain things.
The same issues might be present when, for instance, transparency is heavily emphasized and promoted, e.g. via open practices badges. These open practices badges may however, at one point in time, be changed or turned into something slightly different. For instance, they become a “service” to journals, or somehow monetized, or some “special editors” recommended or even provided by the organization that introduced the open practices badges will start appearing that may influence other things as well.
I think when a person or organization promotes or heavily emphasizes some idea or format (or whatever) it might be important to be aware of the possibilities that this person or organization can benefit from it themselves, or that this person or organization has other plans for things in the future. When this is all known, and purposefully used, by the person or organization I wonder whether this could be seen as some form of grifting or swindling.
“Simonsohn supporting and Nelson and Simmons being silent about their P-curve forensic method and its (continued) use by them and others to cast aspersion on literatures and harm careers…”
Can you give an example of a career being unfairly harmed by the p-curve? It has been established that in extreme edge cases involving contrived data, the p-curve can indeed go off the rails, but based upon those arguments, it seems that the sun will die out before the scenario plays out in real life.
I’m guessing that you don’t have an example, so I would settle for a point-by-point refutation of Simonsohn’s blog post “The p-curve fails if you drop a piano on it.”
[This is where “not a paragon” becomes “not a response.” -ed.]
Quote from above: “Where does this all stand relative to failing to correct or even doubling down on claims whilst representing oneself as Mr Morality / Open Science?”
I came across a recent blog post that appears to be written by Mr. Nosek on the Center for Open Science blog dated july 13th 2026. It reminded me of the quote above, things I have noticed in the last 14 years or so, and things I wrote in my recent manuscript mentioned in another comment here. I wondered about the following, which might be illustrative concerning a few things worth pondering:
This is what can be read in the Center for Open Science blog post:
“The U.S.’s position as a steady foundation for research investment, training, and collaboration globally is lost and replaced with a perception that the U.S. is a tenuous and unreliable partner and funder.”
(…)
“At COS, we are combating retrenchment by centering on our mission and purpose. We are wrestling with fundamental questions about how we can advance our mission to increase transparency, integrity, and trustworthiness of research. What role do we play?”
I thought these sections were noteworthy in relation to what I have written in my recent manuscript:
“Registered Reports are connected to the Center for Open Science (see “competing interests” in Chambers & Mellor, 2018; see “competing interests” in Soderberg et al., 2021), but even scientists associated with such a center might have difficulties with pre-registration. Or how should one interpret that being an employee of the Center for Open Science which offers support to journals, editors and researchers in adopting and conducting
Registered Reports (see “competing interests” in Soderberg et al., 2021), and having published papers about Registered Reports and pre-registration (e.g. Nosek et al., 2019; Nosek & Lakens, 2014), and not optimally adhering to the pre-registration concerning a previous manuscript (see Chatard et al., 2020; Klein et al., 2022, p. 4), may all not be enough to prevent particular pre-registration issues in a now retracted paper (see “retraction note” to Protzko et al., 2024; for further details also see Gelman, 2024)…”
If my section is a fair and correct assessment and description of facts and events, I wonder whether that might point to something important. Perhaps scientists themselves, in this case Mr. Nosek and the Center for Open Science, are partly responsible for things like diminishing trust in science and scientists. And perhaps scientists themselves, in this case Mr. Nosek and the Center for Open Science, should perhaps first get their own sh#t together before writing and talking about “trust” and how “science” is important for X, Y, and Z and such things. At least in my experience, I think his behavior and work, and that of his center, has directly contributed to me losing lots of trust and faith in certain things and people…
As a side note, when reading the blog post by Mr. Nosek on the Center for Open Science blog I was reminded of Bargh’s blogpost about priming effects replicating just fine (also mentioned on this blog here). This is a section of the blogpost by Bargh if I am not mistaken:
“Research has now moved on from the demonstration and replication of priming effects on social judgment and behavior to research on the mechanisms underlying the effects and the moderators, constraints, and limitations of those effects.”
This is what can also be read on the Center for Open Science blog post written by Mr. Nosek if I am not mistaken:
“Previous questions about whether researchers will do open science at all are giving way to metascience questions of whether they are doing it well, and whether open science solutions can meet their promise. Previous questions about whether funders will invest in creating open science infrastructure are giving way to practical questions about how to sustain them.”
As written on this blog here about the Bargh quote about priming effects: “Ummmm, no. Bargh’s research may have moved on, and that’s fine; it’s good to move on and study new things. But for many of the rest of us, no, these effects have not been demonstrated, and the failed replications make the whole thing look like the sort of mess that Paul Meehl wrote about, decades ago.”
Mr. Nosek may have moved on with regard to certain things, and that’s something I think he should perhaps ponder some more. If it’s any help, perhaps it’s okay for me to once again refer to my manuscript that can be found on SSRN titled “Pre-registration, grocery lists, and particular pre-registration issues” concerning this all…
Quote from above: “As a side note, when reading the blog post by Mr. Nosek on the Center for Open Science blog I was reminded of Bargh’s blogpost about priming effects replicating just fine (also mentioned on this blog here).”
I just thought of something else that reminds me of that quote from that Bargh blog post. Here’s another quote from a co-authored paper by Mr. Nosek:
“Despite these limitations, this study provides a basis of evidence to expand the use of RRs in research, and should spur follow-up research to examine the generality of these findings.” (Soderberg et al., 2021, p. 994)”
Is that a particular technique, or tactic? Is that a sub-optimal, or even improper, way to go about things in science?
It looks to me like one can attempt to provide some research as quickly as possible, perhaps you could even be designing and performing that research yourself apparently, for some thing you propose, or want to implement, and are heavily invested in. Then you can subsequently attempt to use that research immediately as a stepping stone to attempt to move to the next phase of a project, while trying to (indirectly) emphasize that these earlier research findings are definitely valid and correct. In that way, you are already halfway to the next stage of things.
I am not sure this is a scientifically optimal, or even proper, way to go about things, as I also note in my manuscript mentioned earlier. The manuscript has several other examples of research where Mr. Nosek has been involved as an author and sometimes even designer of the study if I am not mistaken. When coming across this all, I also started to wonder what the use of conflicts of interest statements even is…
This is depressing. It seems to me that professional association ethical rules often cover this terrain to an extent, but of course these tend to be toothless, rarely interpreted, and there is little effort to educate younger researchers about them. I have been discouraged to hear seasoned and respected researchers tell younger researchers that discovered errors in their work are for others to catch in the “scientific process.” Until the stodgy and archaic approach to academic publishing is supplanted by institutions that do a better job separating process from outcomes, smart research from slop, and that stop reifying “discoveries” as dense little 20-page monographs frozen in carbonite, I predict little change. Academics and publishers seem to have the alacrity and imagination of dead possums (dead *medieval* possums) when it comes to imagining new ways of discerning novelty, talent, value, etc.. I understand that mathematics is currently under upheaval b/c AI is changing the nature of discovery and communication. (Well, Normal Wildberger things so I and I think he makes some good points.) Possibly other fields are going to start convulsing soon as well.
Thank you all for posting. All very informative.
My two cents are that knowingly publishing something that misleads is research misconduct. Failiing to correct any significant misrepresentation, once known, is also misconduct. What surprises me is the extent to which people see the entire process as just a point scoring competition. In that case, published work is a goal — whether right or wrong — and thus somehow inviolate.
I am unfamiliar with the problems with the p-curve analysis. Can someone point me to the reading?
Andy
Quote from above: “I am unfamiliar with the problems with the p-curve analysis. Can someone point me to the reading?
There’s a post on this blog titled “On the poor statistical properties of the P-curve meta-analytic procedure” dated september 25, 2025 which I assume mighth have some more information relevant here.
Quote from above: “What surprises me is the extent to which people see the entire process as just a point scoring competition. In that case, published work is a goal — whether right or wrong — and thus somehow inviolate.”
If I am understanding your comment correctly, that might also be directly or indirectly clear from several arguably problematic issues like using questionable research practices, pleasing the reviewer just to get published, trying to get published in a high-impact factor journal, trying to publish as many papers as possible, exaggerating and hyping in one’s paper to increase the chance of getting a paper accepted for publishing, and largely only submitting significant findings to journals and not really doing something with non-significant findings (e.g. not even mention the non-significant findings in some way in a next paper, or make these available somehow).
That is a really scary story about your colleague. If I did that I would be looking over my shoulder all the time. But maybe that colleague also doesn’t share data.
Elin:
He’s not my colleague any more! I stopped working with him many years ago. He’s had a very successful academic career. He would’ve done just fine without the cheating, but I do think the cheating allowed him raise his career to a higher level. Kind of like Mark McGwire.
I have a funny story about this. My very first published paper in computational linguistics had an interesting idea in it, the so-called “cache language model”, which was designed to improve the performance of automatic speech recognition (ASR) system. At its core, the idea is a very simple one: if you are trying to predict what someone will say next, for instance while dictating a document, you should raise the probability that the next word will be one the speaker has already uttered. For instance, “elephant” is a rare word in English in general, but if the speaker has already uttered that word in the course of the current dictation session, it’s fairly likely he or she will utter it again soon.
My graduate supervisor at McGill didn’t have access to a large-vocabulary ASR system to test my idea, so I tried it out on a large corpus of written documents, using a measure called “perplexity”. Perplexity measures how well your model does at predicting the words in a document: the lower the perplexity, the better the model. I was very happy when my calculations showed that by incorporating my cache idea in the then-standard trigram model for predicting sequences of words, I got a large reduction in perplexity. I don’t have the numbers in front of me, but let’s say that on my test corpus, the cache language model plus the trigram model got a perplexity of 100 as opposed to a perplexity of 300 for the trigram model alone. Clearly, my idea was a major improvement on the state of the art …
Or was it? We announced our results at an ASR conference, and they were published by a major IEEE journal in the field. A few months later, a distinguished German researcher approached me to let me know that he & his graduate students had been unable to replicate my results, on the same data. They did get an improvement in perplexity on the test corpus using my cache idea over using the trigram model alone, but the improvement was much more minor than I’d claimed: let’s say, from 300 to 285 or 280 (again, I don’t remember the exact numbers). I was incredulous, but when I went back over my calculations, I found I’d made a dumb mistake in my perplexity calculation for the cache language model results – forgetting to take the logarithm in one place or something like that. The Germans were right: the impact of my idea on perplexity was much less dramatic than I’d claimed in the IEEE publication.
My supervisor & I did the honorable thing (step 1 in the Gelman ladder): we sent a correction to the IEEE journal in question, thanking the German researcher for pointing out the error. Meanwhile, I was contemplating throwing myself off Mount Royal, or at a minimum, abandoning graduate work altogether. How could my nascent research career ever recover from my first journal article containing an outrageously extreme, false claim? I’d lost all credibility.
Meanwhile, the most prestigious group in the field, the ASR team at IBM Watson, no doubt motivated by the fallacious claim, were trying out my idea in an actual large-vocabulary system – something we’d been unable to do. At the time, it was very hard to improve on the trigram model, which had been developed by the IBM team themselves. Under the right circumstances, they got a very nice improvement in speech recognition accuracy by incorporating my cache idea – let’s say, slightly somewhere in the range 10% – 15% fewer word errors. At a major conference in the field, Dr. Frederick Jelinek, the notoriously difficult and demanding head of the IBM group, went out of his way to credit me and my supervisor for contributing one of the few practically useful new ideas in the field. He tactfully avoided mentioning the massive error in our paper.
That fallacious paper from the early 1990s continues to be cited favorably to this day. In fact, it’s my most-cited paper – and I’ve published a lot of papers since then, some with lots of citations. Nobody cares that the perplexity calculation underlying the results is wildly wrong, because it involves only an indirect measure of the usefulness of the central idea. The central idea works well in practice (not only in ASR systems, but in related fields, such as machine translation), and that’s what matters to most people. For a while I was writing to researchers who cited the paper to start citing the correction, which came out a year later, as well, but I’ve given up (it is cited occasionally).
Obviously, I was damn lucky. I’ve had a good career, because it turned out the central insight of my first research paper was valid even though the paper contained an extremely dumb (though honest) mistake. I wonder if this has happened to anyone else? E.g., a junior biomedical researcher claims dramatic results for curing induced eye cancer in lab mice with a chemical compound he’s devised. A few months later, the poor guy has to confess that his results were erroneous – he screwed up the Excel spreadsheet by mistake – most of the afflicted mice are no better than before. But in the meantime, tests at a major medical facility have shown significant improvements on human cancer patients suffering from a variety of cancers, using the same compound. His career flourishes :-)