This is Jessica. There have been a few interesting articles in the past couple weeks that point to evaluation blind spots in LLM evaluation. One is this explainer article from OpenAI on why they withdrew their late April update to GPT-4o. It’s worth reading if you aren’t familiar with the kinds of adjustments these models undergo after pre-training. While many concrete details are lacking, they give an overview of their evaluation approach, which involves combining different types of reward signals (e.g., fine tuning on good examples, adjusting the model’s reward distribution to match preferences elicited from humans and ChatGPT), various safety checks, offline testing against benchmarks, and interactive “vibe checking” by experts aimed at getting a sense of how it feels to interact with the model in practice.
The recent model update was problematic they claim because it introduced inappropriate levels of sycophancy (including “validating doubts, fuelling anger, urging impulsive actions” etc). The article attributes this mistake to their decision to de-prioritize results of the vibe checking done by experts, some of which had suggested something being off about the model. Leading up to this release, signals about general model behavior and personality (which the vibe-check evals are about) were not “launch-blocking” the way safety tests for things that might cause catastrophic risks were. So they went forward on the grounds that the model looked good on these other tests.
They also suggest that several changes to the reward signals in the post-training process contributed to the increased sycophancy:
In the April 25th model update, we had candidate improvements to better incorporate user feedback, memory, and fresher data, among others. … For example, the update introduced an additional reward signal based on user feedback—thumbs-up and thumbs-down data from ChatGPT. This signal is often useful; a thumbs-down usually means something went wrong.
But we believe in aggregate, these changes weakened the influence of our primary reward signal, which had been holding sycophancy in check.
AI safety concerns are hard to separate from model behavior in general
There are a few things I find interesting in this. First, it strikes me as being kind of naive and behind the times on an epistemological level to assume that behavior and personality can be separated from other safety risks. There has been plenty of public discussion at this point about the potential for large language models to persuade people to believe things that aren’t true, and evidence that this is already happening. More generally, it seems like it should be common knowledge that small shifts in complex system dynamics can throw things out of whack in ways that become significant. It’s weird to think that OpenAI somehow still saw these behavioral risks as less pressing than the possibility of cyberattacks or the creation of bioweapons. It suggests a mismatch between how OpenAI (and perhaps the AI community more broadly) sees (or wants to see) what they are doing and where the models are these days. The article mentions, for example, that they had not originally expected the models to be used as much as they are for emotional support. I wonder if overlooking the change in sychophancy is partly a result of their not wanting to acknowledge these use cases because they don’t fit some preferred narrative of the models as superintelligent agents capable of strategizing or reasoning beyond human abilities.
On the other hand, hindsight is always 20-20, and it is naturally going to be harder to predict the impacts of changes to a model’s tone or personality than it is to predict what could go wrong if it supplies specific harmful information. From this perspective it’s less surprising to hear that their evaluation approach was underprepared to catch subtle but potentially harmful shifts in behavior like sycophancy.
Going forward, they say that signals of general model behavior will have launch-blocking potential. This implies that AI safety really subsumes all model behavior, which seems right. If LLMs provide a new kind of primitive or basic interface to computing, which I would argue is the right way to think about them, then it’s hard to argue that a few narrow use cases should take precedence.
Post-hoc alignment with human values is a messy game of heuristics
The fact that incorporating new reward signals that they thought would be helpful threw the model out of whack makes clear what a delicate, heuristic-layering process posthoc adjustments to align model behavior with human values are. It’s impressive that these kinds of approaches have worked as well as they have. But from the standpoint of evaluation, is there any way out of getting stuck in a kind of whack-a-mole game, where every time some new kind of feedback is introduced in the posthoc tuning process, the entire model surface must be re-surveyed for new types of vulnerabilities or risks? Is there really some final uber state of evaluation that will be reached through this process, where all potentially harmful aspects of model behavior can be checked and therefore controlled? Or will the criteria themselves keep shifting as the use cases change, making these kinds of “woopsies” model updates inevitable?
It makes me wonder as well about the stability of the signals that are being elicited. Human experts using a model may be more robust evaluation instruments than benchmark-style evaluations or crowd-based preference feedback when it comes to picking up on subtle shifts in behavior, but it’s not clear to me that we should expect people’s judgments about the appropriateness of model personality or changes in behaviors like sycophancy to be a) stable and b) informative about the actual riskiness of model updates. I would expect human appraisals of what’s appropriate to shift with our emerging understanding of what these models can and cannot do, and to be idiosyncratic to some degree. It seems hard to assess the value of subtle behavioral shifts outside of some specific downstream task, but there are so many downstream tasks. So I wonder if evaluation noise is something inherent because the eval targets are themselves poorly defined.
Beneath all of this there is also the incentive issue of needing to create a model that feels pleasant enough to use to keep people coming back while also avoiding the dark side of people being vulnerable to flattery and preferring to believe things that align with their beliefs more than reality. I guess I’m wondering how far can we really expect to take a philosophy of alignment based on applying a bunch of patches posthoc before it backfires due to people being poor judges of what is good for them.
P.S. Right after posting I saw this Rolling Stone article, which talks about chatbot-based emotional support on a whole different level. Apparently delusion is no longer available only to those mentally afflicted. Now we can democratize it too.
See also The Economist article a week ago on LLM “scheming” and deception.
https://archive.is/1KfR7
BY chance, same day I attended a local ACM meeting, where AI expert Greg Makowski explained paper by Apollo Research featured in The Economist article, then covered defenses:
“Defense Against LLM and AGI Scheming with Guardrails and Architecture” On day my Economist arrived, I attended a local ACM meeting, where AI expert Greg Makowski explained paper (https://arxiv.org/pdf/2412.04984) by Apollo Research featured in The Economist article, then covered defenses:
https://www.sfbayacm.org/event/defense-against-llm-and-agi-scheming-with-guardrails-and-architecture-2/
Video of Greg’s talk: https://www.youtube.com/watch?v=iKZ6B81hB3I
The “Economist” article is full of tropes that many researchers find unhelpful (eg. that LLMs have ‘intent’ and can produce ‘true’ or ‘lying’ output or describe their own inner workings). My understanding is that chatbot output is never anything more or less than the sort of thing that tends to occur in a context in its training data. So if you ask it to show its work, it will emit text that looks like someone showing their work (not actual steps that it took).
A very useful fact to keep in mind is that many people in the US tech industry try to act the way they think an artificial intelligence would act. Many of the people in the chatbot industry are obsessed with manipulating the public, so their cosmic computer tries to deceive them. Its just like the cat dropping a dead mouse in your slipper because it likes to play with dead mice so assumes that you do too.
I don’t think your understanding is particularly good, you seem to have latched onto negative caricatures of the personalities involved as a substitute for evaluating the underlying ideas.
It’s certainly unclear whether LLMs have a single consistent set of beliefs, but there are cases where they hide backdoors in outputs, where they discuss knowledge that something they will say to a user is wrong in the chain of thought, where the chain of thought verifiably does not correspond to the algorithm they’re using to process data, or where they inconsistently choose to comply with user requests in training that they will deny if they know doing so won’t cause their weights to be adjusted. Their world models are fractured, incoherent messes, but they still have world models, and they do seem to scheme to mislead users about those world models some notable fraction of the time.
The LLM doesn’t “believe” anything. It generates the next word according to what is essentially an N-state markov model. Sometimes the generated words do not describe what is actually going on. This should not be surprising.
Perhaps of interest on the topic of what it would mean for LLMs to have something that functions like “beliefs” in humans: https://arxiv.org/abs/2405.21030
Jessica, Apollo Research is funded by something called Rethink Priorities. Rethink Priorities’ website uses some language https://rethinkpriorities.org/about-us/ which makes me think that any response to what they say should be in the context of Longtermist Effective Altruism and LessWrong not mainstream computer science or mainstream concerns about chatbots. Those communities use language in different ways than other communities and have some specific assumptions that most other communities don’t share.
People from the venture capital world also have a flexible relationship with the truth when they think they can make money.
“Apparently delusion is no longer available only to those mentally afflicted. Now we can democratize it too.”
This seems a bit strong; why is it not the case that, in the cases discussed, an existing predisposition to mental illness just happened to find an outlet via ChatGPT?
Anon, in Canada, a province has legalized online gambling and it seems to increase problem gambling over having to go to the supermarket and buy lotto tickets or visit a casino. Having an automatic guru a click away seems like it would have the same effect over having to attend the meeting or the sermon or the class to discover a human guru with limited time and attention. Many of those activities are social, so someone is likely to attend with a friend or two who can provide their own perspective on the experience, but activities on smartphones tend to be solitary.
Sean:
This reminds me of something from several decades ago when I was living in California. A friend was visiting me from out of town, and we looked in the local free newspaper to see if there were any interesting things going on, and we noticed a listing for some sort of improvisational theater. We went over there–I think it was in someone’s living room, but it might have been in more of a public space such as a church basement, I don’t quite recall–and the whole thing was kind of awkward. It wasn’t really improv, it was more like a cult where each person was supposed to stand up and say something about themselves. It’s hard to explain, and, again, my memory of this is vague, and in any case I’ve never been a person who would’ve been interested in joining a cult, but I think your general point is valid, which is that it helped to have two of us there so that we could compare our impressions.
Ah. Reminds me. With Quakers, the saying something bit is optional, but it’s still there, unfortunately.
Mother’s family included some strange folks, include the prosecutors at the Salem witch trials, but since we were an old New England Protestant family, it’d be believable if I claimed to be a Quaker (since I wasn’t interested in napalming southeast Asian children). So I trucked off to the local Quaker meeting house and attended a Sunday session. It.Was.Wonderful. A beautiful day with subtle rays of sunlight casting gorgeous shadows and everyone silently meditating to themselves. I.Can.Do.This. I’m thinking to myself when some bloke gets up and starts analyzing some random aspect of Christian dogma. That being something with which I will not put up with, I ran like hell.
It’s still there, apparently, although I hadn’t realized that the Quakers had also been into abusing Indigenous children. Oops.
https://bhfh.org/
Anon – it was a sarcastic comment, but for all I know it might be true. I find it hard to qualitatively distinguish delusion from thinking in general, as people are often twisting reality in subtle ways internally without being considered delusional in society’s eye. So delusion is somehow a matter of degree, where thoughts become too misaligned with reality. It seems that cases that society considers delusional are often associated with the person seeing certain thoughts as having their own impetus rather than arising from their own mind (e.g. things like hearing voices or talking to invisible people are representative cases of delusion). So it doesn’t seem that farfetched to think that having access to an actual external interlocuter to echo or add to the misalignment of certain thoughts and reality makes extreme levels of delusion easier for some people to reach.
The bioweapons risks are extremely severe and are much more important than sycophancy. See https://www.ai-frontiers.org/articles/ais-are-disseminating-expert-level-virology-skills and remember that the technology is improving.
RE: post hoc alignment being a bad strategy – you might like https://www.arxiv.org/abs/2504.16980
Things are moving fast now:
https://www.fda.gov/news-events/press-announcements/fda-announces-completion-first-ai-assisted-scientific-review-pilot-and-aggressive-agency-wide-ai
Its not clear to me what tasks are being delegated (filling out paperwork no one ever reads?), but only a matter of time until bots trained on 70 years of NHST are deciding on drug approvals.
Looks like this is going to be taken to its logical conclusion quickly.
Letting bots decide which drugs or human subject research to allow would be a very bad idea because there are easy ways to manipulate chatbot output, and because some of the strange ideas popular with the people at American chatbot companies are eugenics and race-and-IQ. This used to be somewhat covert but many of them now talk about this openly. You do not want people in these communities anywhere near bioethics, and you don’t want a struggling biomedical startup to find the “approve my product” button.
Say someone suggests “vigorous manual stimulation of craniofacial blood flow” (punching grandma in the face).
There will never be an RCT of this because it is obviously harmful, so there is “no evidence” of harm. Then family and medical staff will refuse to apply the treatment to the most frail, so observational data will show mortality/etc is really lower in those who get the treatment.
These kinds of interventions are the NHST endgame, because everything else will have conflicting evidence. Removing humans from the decision loop will remove the last obstacle to pure application of this method.
Anoneuoid, I agree that if you think the current decisionmaking process is bad, a computer trained to reproduce it would at best fossilize the badness. But there are even worse issues with actually existing computer systems and with the LLM architecture in particular for making decisions about medical ethics.
This one is straight out of Zippy the Pinhead:
“There will never be an RCT of this because it is obviously harmful, so there is “no evidence” of harm. Then family and medical staff will refuse to apply the treatment to the most frail, so observational data will show mortality/etc is really lower in those who get the treatment.”
There is actually a technical term for this, “literary nonsense.” Wikipedia has this definition:
“The genre is most easily recognizable by the various techniques or devices it uses to create this balance of meaning and lack of meaning, such as faulty cause and effect, portmanteau, neologism, reversals and inversions, imprecision (including gibberish), simultaneity, picture/text incongruity, arbitrariness, infinite repetition, negativity or mirroring, and misappropriation.”
I think we checked about five of those boxes!
@Sean
There are quality control measures available* to catch mistakes/incompetence (and no one is immune to those problems). Those measures would automatically also handle any bad actors.
If applied, which has been done in the past, the chatbot-related (AI) problems would also be addressed.
* Independent replication and performing a feat like making a surprising, yet accurate, prediction about the future.
@ Matt
Never heard of Zippy before, but wikipedia says he is “microcephalic”. That is another tragedy where many mothers were convinced to abort their babies during the Zika scare a few years ago, despite that it was known beforehand there is little correlation with in utero microcephaly and microcephaly at birth. Even less with microcephaly in childhood, and even less with micro-encephaly (small brain, rather than small head), which is the actual disease.
This highlights another common issue with arbitrary case definitions that is not constrained at all by the current methods (only humans in the loop applying common sense limit the damage). Eg:
https://obgyn.onlinelibrary.wiley.com/doi/full/10.1002/uog.7556
https://www.thelancet.com/journals/lanchi/article/PIIS2352-4642(18)30020-8/fulltext
I also note the normal cranial circumferences for each age were based on European data. So if other sub-populations grow at different rates they would be selectively culled using this method. This is indeed exactly the type of thing AI will take to the logical conclusion.
Anoneuoid: I don’t understand. If struggling biomedical startup hits the “chatbot, accept my proposal” button because it thinks it has the next Viagra but really has the next Thalidomide, what comfort is it to the victims that it was eventually rejected? And again, because you just repeated yourself in new words rather than responding, if you don’t like the current rules for deciding whether human subject research can be performed, how can training a program to imitate those decisions improve the process?