The story
A reader who would like to remain anonymous writes:
A story of research fraud just broke that I’d like to bring your attention to, if you haven’t heard it already.
Last year, a first year phd student in economics named Aidan Toner-Rodgers gained headlines for his paper on AI boosting scientific discovery. Acemoglu calls the work “fantastic” and David Autor was “floored.” I’m sure Autor was floored again when it was discovered that the entire thing was fake.
Hey, that’s a funny line!
My correspondent continues:
It seems TR made up everything; there likely was never an experiment in the first place. He even registered a domain name to look like Corning, who then sent him a cease-and-desist. This blog post by a material scientist points out some of the obvious red flags, as does this twitter thread by a professor of materials chemistry. MIT sent out a press release. They’ve asked arxiv to take down the paper and TR is no longer a student at MIT.
It’s an incredible story, especially given how popular that paper was upon the release of the preprint. It received an R&R from QJE and slipped by Acemoglu and Autor. I’ve been happy to see that the story of the fraud seems to be receiving about as much attention as the paper itself did upon release, but still sad that this could happen in the first place.
Hopefully my next correspondence with you is about research and not whatever this is.
“Whatever this is,” indeed!
I was curious so I googled the usual suspects (*Aidan Toner-Rodgers gladwell*, *Aidan Toner-Rodgers freakonomics*, *Aidan Toner-Rodgers NPR*, etc.), and some fun things came up:
From a Freakonomics episode, Is San Francisco a Failed State? (And Other Questions You Shouldn’t Ask the Mayor):
There was just a study from a grad student at M.I.T. describing how researchers using A.I. were able to discover 44 percent more materials than a randomly assigned group that didn’t have access to the technology.
From NPR’s Planet Money:
One of the big questions in economics right now is which types of workers benefit from the use of AI and which ones don’t. As we’ve covered before in the Planet Money newsletter, some early studies on Generative AI have found that less skilled, lower-performing workers have benefited more than higher skilled, higher-performing workers.
For economists like MIT’s David Autor, these early studies have been exciting. . . . Another recent study by MIT economist Aidan Toner-Rodgers found something similar. It looked at what happened to the productivity of over a thousand scientists at an R&D lab of a large company after they got access to AI. Toner-Rodgers found that “while the bottom third of scientists see little benefit, the output of top researchers nearly doubles.” Again, AI benefits those who can figure out how to use it well, and, it suggests, that in many fields, top performers could become more top performing, thereby increasing inequality.
Tyler Cowen shared the abstract of the Toner-Rodgers papers and none of his commenters sniffed out the problems.
Given that the paper fooled the reporters at the Wall Street Journal and the commenters at Marginal Revolution, it shouldn’t be such a surprise that perennially-credulous outlets such as Freakonomics and NPR fell for it too.
A smooth-looking research article . . .
After doing that quick web search, I followed the second link above to read Toner-Rodgers’s article. It was very well-written: it reads like a real econ paper! No wonder Acemoglu and Autor got conned: the paper is smooth and professional in appearance, down to the footnote on the first page thanking 21 different people as well as “seminar participants at NBER Labor Studies and MIT Applied Micro Lunch for helpful comments.” I’m reminded of the Technical Note at the end of the zombies paper.
. . . with some weird references
The first thing that jumped out at me was this in the reference list:
Diamandis, Peter. 2020. “Materials Science: The Unsung Hero.”
That guy’s a notorious bullshitter–how did that reference get into an otherwise serious-looking paper?
Actually, a lot of the references in Toner-Rodgers’s paper are incomplete, with no publication information at all, just a pile of things pulled off the internet, for example:
Bostock, J. 2022. “A Confused Chemist’s Review of AlphaFold 2.”
Cotra, Ajeya. 2023. “Language models surprised us.”
Ramani, Arjun, and Zhengdong Wang. 2023. “Why Transformative Artificial Intelligence is Really, Really Hard to Achieve.” The Gradient.
Schulman, Carl. 2023. “Intelligence Explosion.”
In the words of the late Joe Biden, “C’mon, man.”
The paper also cites 6 articles by Acemoglu and 6 by Autor . . . ok, I guess that’s why they were “floored” and thought the work was “fantastic.” If only Toner-Rodgers had found his way to including 10 references for each of them, maybe he could’ve moved up to “bowled over” and “amazing.”
Seriously, though, setting aside the junk references, I don’t know that I would’ve noticed any problems with the paper had it been sent to me cold.
Suspiciously wide confidence intervals
The only obviously suspect bits are Figures A.2 and A.3:
The intervals look too wide given how close the points are to the line. But my reaction in seeing something like this is that the model is probably misspecified in some way, or maybe the authors are reporting the results wrong. These graphs don’t scream “Fraud!”; they scream, “Someone is using statistical methods beyond his competence” (as with the multilevel model discussed here).
Funny p-values
Oh, yeah, there’s also Table A1:
Something funny about that first column, no? The estimate is 0.195, the standard error is 0.105, and it’s listed as significant at the 1% level. But 0.195/0.105 = 1.86, which, under the usual calculation, has a p-value of more than 5%.
And Table A3:
0.024/0.015 = 1.6, but a z-score of 1.6 is not significant at the 5% level.
And then these:
First, this looks wrong because with Poisson and negative binomial regressions, you’ll typically get similar point estimates but a wider standard error for the negative binomial. But here the point estimates are much different, and something seems wrong with the standard errors: the negative binomial has tiny standard errors. And again the p-values don’t match the numbers in the table. The estimates in the first two columns of Table A8 are a stunning (one might say, suspicious) 10+ standard errors from zero, but they’re starred as not reaching the 1% level of significance.
It’s kind of amazing for someone to have put so much effort into (allegedly) faking an entire study and then get sloppy at that last bit. Maybe he should’ve faked all the raw data so as to ensure internal consistency of his results. I kinda wonder where all these numbers came from. Maybe he used a chatbot to produce them? It would be kind of exhausting to construct them all from scratch.
That all said, had I been a reviewer I might have pointed out these anomalies, and then the numbers could’ve been cleaned up in the revision process and I’d have been none the wiser.
How did they spot the fraud?
OK, so my next question is, who figured out the paper was a fake, and how did they figure it out? From the Wall Street Journal article:
[Acemoglu and Autor] said they were approached in January by a computer scientist with experience in materials science who questioned how the technology worked, and how a lab that he wasn’t aware of had experienced gains in innovation.
Credit to Acemoglu and Autor for accepting this and not trying to shoot the messenger. Also credit to MIT, which did better than Columbia, UCB, and USC in handling research misconduct. I guess it’s easier to discipline a misbehaving student than a misbehaving professor. In any case, as an MIT alum, I appreciate their statement, “Research integrity at MIT is paramount – it lies at the heart of what we do and is central to MIT’s mission.” In this case, they talk the talk and they walk the walk.
To learn more I followed the link above to the material scientist’s blog post, which gives lots of details on suspicious aspects of the paper, various things that I wouldn’t have noticed–no surprise, given that the last time I published anything in material science was over 40 years ago! I recommend you read the whole thing (the material science post, not my old physics paper).
A story worthy of Borges
Above I asked, how is is that this student reportedly went to the trouble to make up an entire study, complete with a fake webpage, and complete and submit a long, professionally-written research paper based on fake study, and then fall down on his p-value calculations. Converting a z-score to a p-value, that’s the easiest thing in the world, no?
But after reading the report by the material scientist, I’m not so sure, as the p-values appear to be the least of the issues. If anything, Toner-Rodgers should’ve put less effort into faking the statistical summaries and more work into designing a more convincing fake study.
But it’s hard to design a convincing fake study. Fake things look fake. Reality is overdetermined. Remember that Borges story with a map that is on a one-to-one scale with reality? Anything else would be unrealistic. Similarly, if you want to fake a study, it should be coherent, and the only way to do that is to not just fake the tables but to fake the raw data, but then outsiders can check the raw data and find evidence for its construction, so really to be on the safe side you need to actually gather the data, which means you need to perform the study.
In short, the only way to produce a truly convincing fake experiment is . . . to do the experiment for real. But that would take a lot of work! (Also there’s a risk with real data that you might not find the effect you’re looking for, but modern methods of data analysis have enough researcher degrees of freedom that this shouldn’t be a problem.)
So, yeah, Toner-Rodgers showed real talent in writing a real-looking paper with lots of almost-real-looking tables and graphs–but perhaps his most impressive achievement was in “social engineering”: whatever it took for him to persuade Acemoglu, Autor, and others that he’d done a real study. You gotta be a stone-cold faker to pull that off.
A solution that should make everyone happy
The above-linked news article said that the author of this apparently-fraudulent paper is no longer at MIT. But this shouldn’t be a problem. When authors of fraudulent papers leave MIT, they usually can go to Duke, no? There must be a position in the business school for this guy. He has a great future ahead of him. Maybe some Ted talks?
Failing that, I’ve heard that UNR is doing a search for dean of engineering. “Artificial Intelligence, Scientific Discovery, and Product Innovation” sounds perfect for that, no? But really I think that Duke’s Fuqua School would be ideal, a place where he could be mentored by one of their senior faculty with very relevant expertise.






Acemoglu and Autor eventually did the right thing, true. But this story shows how much the top ranks in economics are a circlejerk. Toner-Rogers wrote a paper modeled on many other “cute” papers that praise the Greats. The Greats accepted the praise because it enhanced their reputations, and because Toner-Rogers was a student at MIT and was thus already a junior member of the circle. The QJE refereed a paper without any relevant expertise.
What distinguishes this story from so many others is that the entire thing was made up. Typically only the most important parts are made up, allowing the Greats to dismiss criticism as “interpretation.”
What would make this story complete is if he did the whole thing using AI. Sort of meta then.
The guy’s name is Aiden — “AI” is hidden in plain sight! Not only is the paper a fraud, it’s author is a devious humanoid robot.
AI-den it
Well, the “social engineering” part was not that difficult. IMHO, Acemoglu cherry-picks his methods and evidence to support his conclusions. His “Why Nations Fail” work is full that sort of legerdemain. You can find good write-ups and analysis of this by Jostein Hauge (Cambridge), Yuen Yuen Ang (Johns Hopkins), Peer Vries (U. Vienna), myself (Georgia Tech), and others.
To that list you could add Stephen Broadberry.
When this first came out, I thought, Why did someone who was already at MIT risk throwing away their career by making this up? I subsequently realized that Toner-Rodgers probably got into MIT on the strength of an early draft of this made up paper, i.e. he started in 2023 and the experiment is claimed to have occurred in 2022.
And then I thought that would have been wild if AI could have achieved those sort of results back in 2022. The models I used back then–which seems like an eternity ago–could not do basic algebra using sigma notation without completely hallucinating. So yeah, way to much credulity.
On top of the fact that this is a single author. I mean if you’re a company big enough to employ 1,000 material scientists, why do you give your data to a random first year (or even not yet a Ph.D student) to analyze it? C’mon man indeed.
But at the time it came out, I must admit that I did see the abstract and no alarm bells went off.
It’s easy to understand why a PhD student commits fraud: they need to publish to graduate. doesn’t matter if you’re at MIT, you still need to graduate. And preferably something impressive to get a good postdoc or faculty position later.
And then postdocs commit fraud to get TT positions and TT faculty commit fraud to get tenure. Even someone with tenure might commit fraud to get a better job at another university.
I’ve always been baffled by Marc Hauser. He was a tenured professor at Harvard. Where else did he have to go?
Adede:
I’ve never met Marc Hauser or studied his work carefully. My guess is that he believed his own hype–after all, he was a tenured Harvard professor! He was a certified genius! Noam Chomsky thought the world of him! He could well have been sure that his theories were true, and from that perspective his research assistants’ job was to interpret the data he gathered so as to support his theories. When his assistants refused to do it, they were the bad guys, in his opinion.
From that point of view, Hauser was doing nothing wrong. He was being a good scientist–just like Galileo and other great men of science, he was coming up with theories and then gathering data to support them. I see no reason to think Hauser cheated to advance his career; I’m guessing he did what we would consider to be cheating because he thought that’s how the best science is done.
Let me use your hypothetical reasoning. When a second student remarks an inconsistency in data or a wrong reasoning, a good scientist begins to check. Because that’s science. Doing it differently is bad science.
Raul:
I agree 100% that Hauser did bad science. At least, his experimental work was bad. I can’t comment on his theory. Also, yes, I think he cheated! All I meant with this comment above was that I don’t know that he thought he was cheating.
That said, he was into that “Evilicious” thing so, so maybe he did see himself as some sort of diabolical genius. Who knows? There are lots of stupid people in this world who seem to think they’re geniuses, when all they are is glib and unscrupulous. Glib + unscrupulous is a sort of cheat code, a way for a stupid person to seem like a genius.
Briefly, Toner-Rodgers profile indicates that he worked at the NY Fed for 2 years, as a research analyst. This alone would have afforded him the possibilities of getting to know people who could write letters that might get him into MIT.
It would also potentially give him cover for why people would think it’s plausible for him to get access to this incredible data… Or, better yet, would give him cover if not for him working in macro there. From what I understand of Fed structures’, it would be hard for him to have a chance to focus on AI stuff there, barring some off-chance meeting.
Still, this is a big part of the puzzle. Did he come to MIT with a draft? If so, how did he justify it? Or did he present it as something once a student? Someone’s due diligence failed here, albeit we don’t know how and it would be worth knowing more.
Finally, why? He’s dead to academia and research now, or should be. Was he simply hoping to not get caught until he was high up enough or to have captured enough fruits of his fraud? Like, what was the end goal here? Not get caught ever?
To get back to the point. It’s possible, but not necessary for him to get into MIT thanks to a draft or even just the claim he had the data. But I don’t think it’s necessary.
Pedro:
My general impression is that people keep doing what they’ve been doing before, if has been successful in the past. Without any knowledge of this particular case, I’d just say that if this person had cheated in the past–perhaps made up data or plagiarized for term papers and research reports at his job–then it would be natural for him to continue to cheat, as it worked for him before. You and I might consider this sort of cheating to be “playing with fire” or “leaving hostages to fortune,” but he might just consider this to be standard practice.
Also, I suspect that many cheaters don’t realize that lots of other people don’t cheat. I’m guessing that, if you’re the sort of stone-cold cheaters who can lie to people to their face about these things, you also have something missing inside, and it might be hard for you to understand the normie perspective. Such cheaters might very well think that everybody cheats, in which case all they need to do is to be cleverer than most. There was a Columbo episode that turned on this idea, actually!
Andrew:
Yes, both are plausible explanation, both for this case and more generally. Econ and social sciences in general tend to find evidence of habit formation and experience based expectations and that fits the bill quite well. I think these has a good chance of being the truth or part of it.
Still. Assuming the paper wasn’t why he got accepted into MIT, then he was more or less set for life if he could finish the PhD programme*. It’s one of the most prestigious econ programmes in the world and even people who don’t land the best jobs from it tend to get paid very well and can be considered set for life, more or less. The risk of being caught seems disproportionately high for the potential benefits.
Ergo, he may also have miscalculated, thinking the paper would get some attention but not that much (it made big splash!), leading to more in depth analysis than he expected and getting caught.
Today’s blog begins with
“If only Arxiv required researchers to sign at the top rather than the bottom of the page, none of this would’ve happened.”
This, to the cognescenti, is an obvious reference to one of the withdrawn papers of Francesca Gino. Even if everything were legitimate about that study involving Harvard students–and, notice the subjunctive–would we really believe it generalized to Americans, people in general? Far more chicanery is going on in photos known as “western blots” and that needs more coverage because cancer funding is involved. I would tell you more about the subject, but the only thing I know for sure is that “western” has a lower case “w”, unlike the Southern blot which has an upper case “S.”
This is because the western blot is named after the fact that it is traditionally placed on the left side of the figure (the western point on a compass rose with North on top), whereas the Southern blot is placed on the top but named after its inventor Bill Southern. I know this to be true because Chat GPT told me.
Because of the person named Southern, it had an upper case due to being eponymous. Then because of geographical punning, other directions came into the (insider) act. The inventor of the western blot, W. Neal Burnette, has a first name beginning with the letter “W” but I have never been able to find out what that is. His original paper was rejected at first:
https://www.biotechniques.com/biochemistry/south-north-east-and-west-ern-the-story-of-how-the-western-blot-came-into-being/
“When Burnette submitted his methodology to Analytical Biochemistry in 1979, it was not very well received and his manuscript was rejected. This was largely due to the distaste for the naming convention.”
“Since then, other blotting methods have also been invented, including those aptly labeled as the eastern and southwestern blots.”
Other fun and games are done with western blotting, especially data forgery because photo shopping makes it possible to substitute some (possibly rotated) unconnected blot with another. Elisabeth Bik and Sholto David are the two prominent data sleuths who have uncovered massive, outright fraud.
Researchers can just keep doing western blots until they get a “representative” picture they like anyway. Of course this is easier for big labs with lots of funding.
So being able to photoshop what you want is more like an equalizer.
The root problem is there is no expectation to share all the attempts with the reader, which renders the pictures worthless to begin with. Even when quantifying, it is quite normal to “throw away” 1/3 or even more attempts due to artifacts.
Similar for representative histology images. They are more to show a person actually did do something and aren’t really meant to be taken seriously.
Note that my entire post was tongue in cheek, but the Southern blot is actually named after Edwin Southern according to Wikipedia, which is about the most reliable info on the internet, and that seems ridiculously ironic given how much FUD there was when it first came out.
FWIW, there was a paper in Science last year by some Materials Science blokes that used “AI” in it’s title and abstract.
Since Mat.Sci. was my undergrad minor (and my (first try at) grad school major (in which I bit the dust something fierce (MIT grad school is not for the faint of heart or the merely interested))), I read the paper. They were doing some hairy extremely high-dimensional analysis of multi-component systems. In my day, mixing two metals in various amounts and seeing what would happen was the best we could do, so this was a kewl paper. In the extreme. But AI? Not so much. Just gradient descent in enormous spaces combined with some non-local search, some real data, and some domain knowledge. Not just good Mat.Sci., but good Comp.Sci. stuff as well, but since everything’s got to be AI these days, you have to read the papers to figure out what was going on.
Anyway, my first guess at this disaster was that the twat went after the Mat.Sci. types asking questions that assumed “AI” meant “LLM”, and the Mat.Sci. guys, in response to questions that made no sense, gave answers that made no sense.
But it was problably just fraud, not stupidity.
A few years ago I was on a training course where they were showing clips of Wansink talking about soup bowls and the like. Thanks to this blog I was able to point out to the people running the course that his work was completely discredited. For all I know they might still be using those clips. But the point is that regardless of this paper being discredited, there will probably be people using it to promote some sort of innovation tool or strategy – based on the principle that a lot of victims will not look into the work they are referencing.
Pessimistic explanation: he was taken down quick because he was a graduate student. If he was a professor, he would have had the same fate as e.g. the “Why We Sleep” guy (still tenured iirc).
I 100% agree with the sentiment. But what’s wrong with Why We Sleep guy?
https://statmodeling.stat.columbia.edu/2020/03/24/why-we-sleep-a-tale-of-institutional-failure/
I just saw this headline on NPR website, How an AI-generated summer reading list got published in major newspapers. https://www.npr.org/2025/05/20/nx-s1-5405022/fake-summer-reading-list-ai
There are other papers which have had a significant impact and that may deserve careful examination. One is described here: https://theintercept.com/2023/07/12/covid-documents-house-republicans/
The other is examined here: https://weirdtech.com/sci/expe.html. .
Together they illustrate a pattern of wilful dara misinterpretation at the highest level. So high, apparently, that it is safer to ignore it.
Sidenote, The Intercept- and especially Ryan Grim – should never ever be taken seriously though
The article is quite factual. It is referenced in a recent NYT article, where I picked it up https://www.nytimes.com/2025/05/21/opinion/covid-lab-leak.html . Beside, personal innuendos stink.
Ryan Grim was one of the lefty ends of the red-brown horseshoe alliance that was con artist Tara Reade’s PR club. Rich McHugh of Business Insider was the Republican operative who helped them publicize their lies:
https://medium.com/@macarthur.cliff/the-tara-reade-case-eight-things-the-media-wont-tell-you-27d3ca14978
The “late” Joe Biden is still alive!
lol exactly what I thought when I saw it!
..but he was ‘late for dinner’.
Modern Academic Publishing is flawed, journals are flooded with nonsensical research because in many places academics are measured on the number of publications independently of the quality.
Academic journals are no longer a place for discussion and exchange of ideas. They have become a “form” excersie rather than an intellectual exchange, as long as the paper looks well written it might get published.
We are in need of new and better ways to advance science.
Hi Andrew!
I was an exchange student at Fuqua SoB and at the School of Law just a couple and a half decades ago. Coming from Argentina loved the place.
I read the joke about Duke and couldn’t get it. Would you please so kind as to explain or post a link about the situations at Duke that gave the university such a bad name?
Thanks a zillion!
It’s a reference to Dan Ariely.
https://en.wikipedia.org/wiki/Dan_Ariely#:~:text=Ariely%20taught%20at,%5B3%5D
https://www.forbes.com/sites/christianmiller/2021/08/30/an-influential-study-of-dishonesty-was-dishonest/
I’m using this as an example in a and the Toner-Rodgers paper has 93 references in Google Scholar, including 10 as of 2/23 in 2026, and virtually all treating in as legitimate research.
Michael:
No problem: if they do enough bad research and it gets cited enough, they can get an advice column in the Wall Street Journal and have their work be the basis of a network TV show. Also tons of NPR exposure!