Peter Dorman writes:
In case you haven’t seen it, check out this recent piece in Rolling Stone. A key paragraph toward the end:
Craig Callender, a philosophy professor at the University of California San Diego and president of the Philosophy of Science Association, agrees with that assessment, observing that “the appearance of legitimacy to non-existent journals is like the logical end product of existing trends.” There are already journals, he explains, that accept spurious articles for profit, or biased ghost-written research meant to benefit the industry that produced it. “The ‘swamp’ in scientific publishing is growing,” he says. “Many practices make existing journals [or] articles that aren’t legitimate look legitimate. So the next step to non-existent journals is horrifying but not too surprising.”
The one point I [Peter] would add is that the first step to perdition is citing sources you haven’t actually read or at least looked through yourself, relying instead on what other people said about them, or simply that they were used as citations in other articles. I know I’ve been tempted to do this sometimes because it takes extra time to do it right, and I may be in a hurry, but then I manage to impose a little integrity on myself. After you’ve taken the first step toward “blind” citation, however, everything else becomes possible — and now even likely.
This reminds me of something that my adviser once told me, that he made a principle of never putting his name on a document that he hadn’t read.
As to citing sources, yeah, I agree with Dorman 100%. You should only cite sources that you, or one of your coauthors, have read or have at least looked through.
A related point: When you’re reading a paper and going through its references, you will sometimes find that the paper described the reference inaccurately. As we wrote in Section 3 of this paper:
Unreplicable claims based on weak theory can gain apparent support by connections to related published work. Three problems can arise.
First, the connections between the cited literature and the new study can be tenuous, and this can particularly be an issue when the underlying theory is vague. Ideas such as embodied cognition, evolutionary psychology, nudging, mindfulness, or mind-body unity are general enough to encompass a wide range of potential phenomena, to the extent that there is almost no limit to the past studies that could be thought to have some possible relevance to any new experiment.
Second, informal literature reviews are subject to selection bias. An article promoting a controversial idea can easily cite studies claiming to have found evidence for related ideas, while avoiding citations of failed replications or papers suggesting alternative theories. This can even be a problem with systematic meta-analyses, if the entire subfield being meta-analyzed is full of studies with uncontrolled researcher degrees of freedom.
Third, the interpretation of individual studies being cited can be seriously flawed. This is a problem of citing past literature as support for a general claim without looking at exactly what was done in the cited research and without following up on that work. Here we discuss three different examples of this sort of misinterpretation of the literature cited in the paper under discussion. . . .
So, yeah, don’t trust what someone writes about a cited study. Take a look at it yourself.
P.S. I wrote this post several months ago. It’s just a coincidence that it happened to show up the same week as What do I think about that proposed Arxiv policy to ban authors of papers with AI slop?.
Someone quoted in the article says, “LLMs are structurally indifferent to truth.” I think a way to put this that’s easier to understand is that LLMs don’t distinguish fiction from nonfiction. If I ask for an account of what happened, I’m likely to get an account of what probably happened (given what the database knows about the people and that sort of situation), or how a typical situation like that would play out, with some invented details to make it interesting.
It is very easy to make the chatbot produce invented citations, often to real texts, for things “everyone knows”, or things that fit generalizations or stereotypes like “things AOC would say” (or “what did Gelman say about Singal’s chapter on power pose”) and very difficult to get it to tell you whether there really is support for whatever it is instead. It’s also very difficult to get a definite negative “there’s no source for this thing some people have said is so”, I’ve found it will always come back to lying that a source says something it doesn’t say. That can be instructive, but only in a negative way.
I agree, but it is unfortunately true that many people find themselves unable to distinguish fiction from nonfiction. We currently have some Republicans in Congress still saying that Jan 6 was staged by Democrats. And we have a President who repeatedly states falsehoods – whether or not he believes these, I don’t know, but there are plenty of people who accept whatever he says as true. I think we all still believe that humans are capable of distinguishing fact from fiction, but I’d also say that LLMs are similarly capable – given the right prompting. Is that really different than humans who also seem to require prompting to make the distinction?
Sadly, I have to say there may not be much difference between LLMs and humans in this regard. I’d guess that a carefully designed empirical study would find LLMs commit such fallacies more frequently or easily than humans – but that doesn’t provide much solace for me.
I see this kind of equivocation between LLMs errors and human errors all the time and I really don’t see the purpose in it. It is categorically different.
If you label articles explicitly as false and feed them to the LLM during supervised fine-tuning, the LLM will still repeat the claim as if true.
https://arxiv.org/html/2605.13829v1
To this day, SOTA LLMs will still say that 5.11 > 5.9. In the best cases, with the right settings, they will then explain their reasoning and contradict themselves without acknowledgement.
https://media.licdn.com/dms/image/v2/D4E22AQFz1E3-FpXKnA/feedshare-shrink_800/B4EZgsTV3xGwAk-/0/1753089925186?e=2147483647&v=beta&t=F-UvIPdIWxBcutr5OA2N8dHi27iqLqIPo0uVEmVP3hw
The situation is simply NOT “LLMs hallucinate sometimes and humans make mistakes sometimes, it’s just a difference of degree.” No, the way LLMs work is completely different from the way humans think. For many tasks they are obviously incredibly superhuman (if only because of the breadth of ingested material and speed of work) and for others they are obviously broken machines. This kind of equivocation can only impair one’s judgement of what’s going on.
On one level, a machine and a human are clearly different and it is foolish to think they work the same. But when you state that LLMs are “completely different from the way humans think” you must know a lot more than I do about how humans think (entirely possible – but from my knowledge I can’t claim to know how humans think at all). Drawing potential parallels between LLM and human mistakes can impair judgement as you suggest – but it can also serve the purpose of protecting against the all-too-common critique of AI based on one-sided reasoning. That reasoning only sees errors that AI makes while ignoring human errors. That can also impair judgement.
I do not intend my comment to suggest that AI can be trusted or that AI is not dangerous (in many ways). We can probably better understand the ways that AI makes mistakes than the people do – that is about the only hope I have for managing AI to reduce its errors (leaving the bigger danger, in my mind, that AI will work as intended by nefarious actors). I’ve pretty much given up hope that we can understand people well enough to reduce our errors.
Dale Lehman, if you don’t know anything about how humans think, you could go learn. That is one wonderful thing which humans can do and LLMs cannot once their training is finished. A full understanding of the human mind does not exist, but what we do know is very different from how transformer-powered LLMs work.
I am an amazing electro-chemical system descended from hundreds of millions of years of evolution which creates a lot of illusions in itself because those were useful for social animals. A LLM is a remarkable digital system descended from a few decades of mathematical research which can sometimes create the illusion of being the first illusion if it is built the right way and the observer wants to believe. The little bit I know about them is fascinating like learning how magicians do their tricks, or all the tricks for binary math and floating-point arithmetic I learned in my days as a computer scientist (everyone here knows that computers have trouble representing 0.1 because its not a sum of binary fractions and have to perform painstaking mathematical maneuvers to hide that from us right?)
“you must know a lot more than I do about how humans think (entirely possible – but from my knowledge I can’t claim to know how humans think at all). ”
So go read the literature. Philosophy, cog. sci., psychology. Lots of smart folks have thought a lot about how people think, and have noticed that they actual do think and reason. Really, people do think and reason. LLMs don’t. It’s a humongous difference.
There’s lots of work on it. In particular, the basic idea of how LLMs work (statistics on words) has been convincingly debunked as a theory of mind. Read Jerry Fodor, for example. (In Critical Condition, I think mentions this. It’s good reading even if that point was somewhere else…)
For those of you inviting me to learn about how the human brain works, I both defer to your superior understanding and stand by my statements about how little we understand. Here is a high level explanation of how the brain works: https://www.ncbi.nlm.nih.gov/books/NBK279302/. I don’t claim this is the extent of our knowledge – I know a lot of serious work has been, and continues to be, done understanding this. But I don’t believe this “understanding” really explains how we reason. I don’t claim that computers reason the same way that people do, nor do I claim that computers “reason” at all. But I’ll stand by my claim that we really don’t know enough about how humans reason to meaningfully describe what is different about the reasoning that AI displays. The key word is “meaningfully.” Of course, a machine and an organic body part work differently, but exactly how “reasoning” separates one from the other still seems elusive.
I’m willing to become educated if you can point me to sources that provide such explanations, hopefully without my needing to obtain advanced degrees in neurology, biochemistry, etc. I’m just looking for a high level explanation about how human reasoning and AI are different categories that render comparing them useless.
I’m sure there are some humans, who when say given an assignment to write a true narrative, will embellish and enhance their narrative with things most readers associate with “a well-written story,” I guess there are probably some humans who, when told that they’re supposed to write a *true* story, don’t understand what’s wrong with what they wrote, and long experience online proves that there are people (it’s impossible to say how many are trolls just playing a role but presumably a few are sincere, whatever that means in this instance) who, if I say “that source doesn’t say what you say it does, please provide me a quote or say the source doesn’t say that,” will possibly say “I’m sorry” but will continue to insist the source says what I asked about. We don’t consider such responses to be in the normal range of human response, we associate them with children and people with cognitive or emotional disorders. To say that’s desired behavior from an AI rests on an enormous equivocation.
But LLMs are, as the article says “structurally”, built in such a way that they don’t recognize the difference between “tell me what you conclude, based on written sources, did actually happen, and write it up in narrative form” and “write a story based on the idea that this happened.” I understand the firms try to train the AI to do this less, and it isn’t a “computer” problem but a specifically “LLM” problem, but it’s a fundamental problem.
I have two points of disagreement with you. First, your notion of “the normal range of human response” is a dangerous fiction in my view. It is more of a normative statement about human behavior than an empirically based one. It is socially constructed, thereby running the risk of describing how you would like people to behave rather than the way they actually behave.
The second is the belief that AIs, by construction, cannot be trained to recognize various differences. I am far from sure what the limits on training are. My experience thus far, is that LLMs struggle with context – they often don’t recognize important features that should be taken into account. Unfortunately, I share that problem. But in areas where I know enough, I can recognize these deficiencies and correct for them. So I see human education as essential for use of AI – not just education of how to use AI, but enough subject matter knowledge to be able to recognize AI errors. If your meaning is that AIs, by construction, cannot be trained to be completely autonomous then I might agree. But more importantly, I have no interest in completely autonomous AI.
Dale,
I suspect we have basic philosophical disagreements and should end the discussion here. But I notice you’re not addressing the question why we would want an AI which in this case is the question whether knowing the limitations the kind of AI people are using for literature searches now would help people do their research better.
“why we would want an AI which in this case is the question whether knowing the limitations the kind of AI people are using for literature searches now would help people do their research better”
I can’t answer that question because I can’t understand it. In any case, I’m not sure why you are seeing anything in what I am saying as my “wanting” an AI for anything. I find it fascinating and useful personally, but I feel quite sure that it will lead to all sorts of problems that we humans are quite equipped to misuse and ill-equipped to deal with.
Just have the AI return an exact quote and check it yourself. Institutionalizing this practice would be a step up from the current situation where people cite an entire paper/book with a single number like [1].
It was already extremely common for the references to not contain the referenced info before the bots.
“Institutionalizing this practice would be a step up from the current situation where people cite an entire paper/book with a single number like [1].”
As I’ve argued elsewhere, requiring page numbers by default for each citation would be a good start. I try to do that more now myself to encourage the practice.
The topic is also related to the folklore claims online that “academia still uses the arcane PDF format”. PDFs are great because they can be stored locally. Indeed, I have a local copy of almost everything I cite. That practice too should be encouraged. Someone could just ask: “we have some doubts, so please send the associated references.tar.gz” file for us to see.
Otherwise, why are we even having this discussion? Has science fallen so down that we need to return to the ABC of it?
@Jukka
Page numbers would be an improvement but as someone who does follow reference chains I would much prefer an actual quote or figure/table ref.
I often find claims like “x is expressed by 3T3 cells [21]” where the claimed info is not in ref 21. Maybe there is speculation about the topic in the intro or discussion, maybe ref 21 cites another paper that putatively includes the info, sometimes its not present at all.
This is a huge time waster because it is much harder to verify something is NOT present rather than present.
I actually prefer the llm method of just making up a source to the human method of using a real source that does not contain the info. But best is to just include a direct quote, space is not an issue in the digital age.
Maybe one way to think about it is that LLMs are trained on a wide range of human-produced texts, including many that are dishonest or at least tendentious. What’s “right” to them is what is consistent (at a rather deep and incomprehensible level) with their training. If this is true, we could conjecture that, if they could just be trained on the output of some superior, always trustworthy species, they would be trustworthy themselves.
Yes, it is very basic scholarship that you should never cite something as source for a fact which you have not seen yourself, and should be very careful about citing something as a resource which you have not seen. It does not matter whether you get the note from a LLM or an encyclopedia, good scholars check what the source actually says.
I think the response to this shows just how precarious life in lab science has become, because in a healthy workplace, people trust each other and have systems of verification wherever errors are frequent or dangerous. But in a lab where postdocs and students are coming and going every year or two, it is hard to build that kind of solidarity and create those systems.
Anne-Wil Harzing has a classic paper on the problem of empty citations in business schools (“Are Our Referencing Errors Undermining Our Scholarship and Credibility? The Case of Expatriate Failure Rates”).
There’s a standard way of citing sources you have not seen, but instead have relied upon an account of what the source contained. Here’s an example, taken from Simmonds, N.W. 1960. Flower colour in Lochnera rosea. Heredity 14:253-261.
FURUSATO, K. 1940. Polyploid plants produced by Colchicine. Bot.& Zool., 8,1303-1311. [Not seen: cited by Darlington, C.D., and Wylie, A.P. 1955. Chromosome Atlas of Flowering Plants. London.]
Don’t just skim or read sources, but appraise them. At least, that’s the lesson that I learned from this blog.