This is Jessica. I don’t usually post multiple times a week, but turns out I have more to say on the topic of machine learning and human problems.
“Alignment” is the term used in the AI and ML communities to refer to the goal of aligning machine learning models with human values and preferences, so as to avoid risks ranging from the mundane to the catastrophic. It’s the topic of papers, workshops, talks, funding calls, etc.
There has been criticism of the nebulousness of what alignment is supposed to actually represent. Some of the critique of the ML conception of alignment comes from HCI research, the very interdisciplinary field that studies how people interact with technology and how to design human-computer interfaces. This pushback predates the “alignment” buzzword actually. I remember watching many in the HCI community bristle when in 2016 Michael Jordan wrote a blog post calling for the creation of a new “human-centric engineering discipline.” Seeing human-centered concerns get called out as a new frontier in AI/ML circles was enough to motivate some of the better resourced HCI researchers to create centers on Human-Centered AI or install themselves as team leaders in big tech companies, ensuring they wouldn’t be overlooked. Others have worked to make AI-related applications a bigger part of HCI research. Many are left to stew about wheels being reinvented, trying to be patient and issuing the occasional plea for everyone to recognize the overlapping goals.
My take is that HCI can help quite a bit with alignment, but that what it can offer is not what much of the ML research community wants or perceives themselves to need. It’s kind of like what a consulting statistician can offer to a data analysis versus what they are perceived to offer by those that recruit them. The real value of adding the statistician is often their role in helping you rethink your objective from the ground up. It’s not necessarily that they’re going to give you exactly the best tools to address some narrower problem you’ve convinced yourself needs to be solved. E.g., you’re convinced that if you just find the right causal inference technique you can confound a big messy dataset you’ve amassed and learn exactly how to improve some outcome X, but the pesky statistician comes in and spoils it by telling you, “No, if X is ultimately the goal, you’re going to need a different data collection procedure altogether.”
In the case of aligning ML, there are certainly human-oriented questions that arise within the current paradigm for aligning models. For example, questions of eliciting specific information from humans become important for deploying generative models. Reinforcement learning from human feedback (RLHF) is a standard method for fine-tuning a large pretrained model like GPT-4, where some group of annotators is recruited, often with no special experience required, and asked to select their preferred model output in a series of forced choice tasks, usually given some loosely defined criteria like “most helpful” or “least harmful”. Behavioral models for aggregating preferences across people like Bradley-Terry-Luce are used to learn a utility function. Human-oriented concerns include how to design the forced choice task and interface, how much information can reasonably be obtained from a single person, and how to crowdsource this efficiently. Beyond the common need to collect human annotations, other examples where human concerns arise in the current ML paradigm include questions like how to represent fairness ideals or how to evaluate post-hoc explanation techniques.
Could an HCI researcher be helpful for these questions? Sure, though I suspect that the most relevant work for some of these elicitation problems is likely to be found elsewhere, like psychophysics or decision science. Could the ML researcher figure this kind of stuff out without the HCI researcher? Probably. In many cases it may be more efficient for them to do it themselves, since HCI is a very large and interdisciplinary field. So I’m not surprised that ML researchers are often doing these things themselves, nor do I blame them.
On the other hand, I think the HCI pushback to ML alignment is valid when you consider the broader goal of creating predictive models that are well-aligned with human goals and values. If there’s a secret sauce that your average HCI researcher can bring, it’s the mindset of user-centered design, which makes serious attempts to understand the needs of the people being designed for. HCI research also demonstrates what it looks like to hold the conviction that human values are not monolithic, contributing knowledge on a variety of methods to try to get at what different groups want from technology. When taken to heart, I expect this kind of perspective suggests rethinking pretty much everything about human-facing ML models from the ground up.
Unfortunately though, I don’t really see much incentive for the average ML researcher interested in alignment to invest in the HCI way of doing things. All interdisciplinary collaborations tend to be hard, and this one seems likely to be particularly slow and messy. Meanwhile AI/ML research is moving at a faster pace than ever.
I also tend to believe that when someone is peering into a field, and believes that they can bring in some new perspective or methods that will be transformative, there’s an onus on that person to invest enough in understanding the field they hope to change to be able to demonstrate the value they want to bring. You can’t really expect people to listen if you haven’t taken the time to understand their concerns well enough to show them that you really could provide concrete suggestions. If the HCI researcher wants alignment to be done better or differently, maybe it’s time they temporally reinvent themselves as an ML researcher. Figuring out how to publish HCI-oriented papers at ML venues may not be easy, but it’s a step toward real impact.
I don’t mean this last part to sound dismissive, or like I’m trying to defend ML alignment. I think it’s just how things work. I’ve had multiple times in my career where I’ve looked at some other field and thought, I bet I could improve that. It’s how I’m feeling right now, actually. And every time I’ve been in this position, it seems clear to me that the only way to have that impact is to invest enough time in the new field to internalize how they think about it.
My general rule is to never trust an academic outside their narrow field of expertise and to realize within that narrow realm they are usually heavily invested in a particular point of view. Coming from CS, it surprised me how cautious the applied math and statistics folks are compared to computer scientists at generalizing beyond their training.
If you think ML academic appropriation is bad in HCI, just consider stats and linguistics!
OpenAI crowdsourced their human feedback. There are details in the GPT-3 paper. They had raters rank the outputs in order of preference and also rate them on a 1–5 scale for quality. They also developed example question/answer pairs for training (can’t recall if that was also crowdsourced, but probably not). What I’m most surprised about is how little money they spent on human feedback compared to, say, electricity to fit the foundation model. You can see from the llama fine tunes like starling that GPT’s “personality” is entirely driven by the fine tuning.
The elephant in the room here is the lack of shared “human values.” Cultural consensus theory is an interesting take on this that’s related to crowdsourcing—it’s just the David and Skene crowdsourcing model with the truth represented as a mixture rather than a single value. One way to see that values aren’t shared is to look at all the criticism of the values employed by OpenAI to guide ChatGPT’s alignment from the right. You can read quite a bit about how OpenAI chose crowdsourced workers who themselves passed an alignment test in the GPT 3 alignment paper.
I hardly see how psychophysics is relevant, but psycholinguistics is. I don’t know what decision science is and a quick web search didn’t help.
With GPT you have a person interacting with a computer, so it seems like HCI would be relevant unless the field’s painted itself into some kind of narrow corner (CHI always did seem very stylized to me). For example, how to run a focus group is rather relevant here! Right now, the OpenAI UX is super simple—pretty much still just a chat interface, but presumably that’s not true of all the applications using LLMs, like CoPilot.
P.S. Why do you call “alignment” a buzzword? It seems like a reasonable choice of technical term to me. Is there a term you’d prefer? Every field develops specialized terminology (e.g., “affordance” in HCI/UX), which outsider always call “jargon,” just so people don’t have to use long descriptive sentences every time they want to bring up a concept. “Alignment” is jargon for NLP the same way “home run” is jargon for baseball and “stop sign” is jargon for driving and “fork” is jargon for a dining tool. That’s just how language evolves—we invent things and have to give them names so we can talk about them.
Yes, it’s naive to act as if there are universally shared human values. That’s one of the things I’m pointing out that HCI is good at accepting and trying to account for.
>I hardly see how psychophysics is relevant
You have a bunch of ML people using forced choice-style tasks to elicit human input, where people are given some criteria and asked to rank stimuli. Questions like how many items to give them and how that affects noise rates, how to model stupid errors they make (lapse rates), what kind of crossover effects you might expect over repeated trials, etc become relevant.
Nowhere am I saying I have a problem with the term alignment. But given that it’s the fashionable term to use when you want to nod to lofty goals related to taking human needs and values into account, steering clear of existential risk, etc, I called it a buzzword.
“Alignment” is the property that distinguishes an AI that can make an exact duplicate of a strawberry on the cellular level from one that can do so and also won’t destroy the world as a side effect.
It has been said that “AI alignment” is associated with the LessWrong and Effective Altruism movements and mainstream research uses a different term. Jessica links to a post by OpenAI which had several members of those movements on its board last winter.
What’s the different term?
Possibly AI safety? A detailed sketch of these movements, their connections, and what topics you should be wary of them on is “Extropia’s Children” https://aiascendant.substack.com/p/extropias-children-chapter-1-the-wunderkind
Thanks for that link, Sean!
By the way, the second “chapter” of that article is seriously fun.
https://aiascendant.substack.com/p/extropias-children-chapter-2-demon-haunted-world
And the fun continues on in later chapters.
David, you are welcome! There was a lot of shouting about those spaces last year but that series is detailed without being false or nasty.
I also want to say that I am not an academic computer scientist like Jessica, just someone who has observed that LessWrong and associated movements generally use language differently than most academics do.
I’m not sure how serious you are about wanting to reinvent yourself as an ML researcher, but the amazing folks at EleutherAI (eleuther.ai) have a lot of experience training, evaluating, and finetuning language models, but absolutely lack the expertise and perspectives you discuss here. I’m sure they would jump at the chance to collaborate.
I’m pretty serious… spending much of my time these days reading ML papers. Feel free to send me an email.
I think HCI is actually at the crux of many of the recent advances in AI. The chat interface of ChatGPT was hugely important for the success of that product, and still more work needs to be done to understand how humans are going to interact with these systems. The chat interface probably won’t last forever-the way we interact with these models, and others in vision and RL, is yet to be developed!
Looking forward to seeing what you do in this space :D
The way you are thinking about AI alignment (or safety) is very similar to my conclusions. People who are good at designing AI models and training regimes are not necessarily good at thinking about how the models interact with societies and what are some of the problems that might arise. Some of the pretty critical issues are often very poorly conceptualized, often caused by the fact that the data collection and evaluation process is designed by an AI engineer without any input from experts from various relevant fields. This is also related to how AI as a community put huge emphasis on how well you can develop AI models, in comparison with the relatively small emphasis on how well you can design a data collection process (in short, the perception is that model work > data work).
For example, consider the case of gender bias. Even the LLMs developed by multi-billion dollar corporations are often only evaluated with datasets such as WinoBias. This dataset measures how much does a model associate certain occupations with he/she pronouns (scientist is a he, nurse is a she, …). This is certainly “a” gender bias, but do they really believe that this exhausts all the risks related to gender inequality in LLMs? Similarly many gender bias datasets are a crowd-sourced set of samples, but nobody is actually analyzing what is the composition of gender biases in the collected data, and which biases were left out completely because the crowd-source methodology simply forgot about them. Other societal biases and issues have similar conceptualization problems, and in effect, we just do not know about many of the risks these models have because the expertise is not there and people in charge are instead messing around with different training algorithms, because that’s what AI scientists are supposed to do!
However, I am not sure whether HCI is the solution. I think it is probably “part” of the solution, but it is probably not enough by itself. I think that it is necessary to consider what experts from individual fields (gender experts, disinformation experts, education experts) are saying about the LLMs and take that into consideration. AI as a field is simply not strong enough to cover all these bases.
I agree, it’s not going to be the whole solution.
Quote from the HCI paper:
“For over 40 years, HCC has struggled with, and made progress on, “aligning” different technologies
to people. Key issues include: Who, exactly, are we talking about (e.g., [1, 4, 5, 11, 24, 38, 42])?
How do we know what they want (e.g., [8, 13, 16, 21, 26, 37, 41])? How stable is what they want to
do with technology (e.g., [39])? Can they help us design it (e.g., [19, 20, 25, 31, 34])? What are the
different ways a technology can be designed (e.g., [6, 33])? How do we know if it’s good for people
(e.g., [9, 32])? What are the limits of design- and tech-centric approaches (e.g., [15, 18, 35, 43])? Is it
possible to avoid baking systemic oppression into technology (e.g., [3, 7, 10, 14, 22, 27, 29, 36, 40])?
Casting alignment as HCC invites it to draw upon the considerable theories, methods, and findings
of HCC, and the fields from which it borrows (e.g., STS, Communication, Ethics, etc.)—instead of
re-inventing them under new names. Perhaps we don’t need the word “alignment” at all.”
I buy the idea that HCI has valuable things to say about aligning AI systems. But it’s worth noting that what’s referred to as “alignment research” seems considerably broader than HCI. Check https://www.alignmentforum.org/ for a feed of the latest research. Search for “overview” in this post to get a list of research overviews: https://www.alignmentforum.org/posts/QBAjndPuFbhEXKcCr/my-understanding-of-what-everyone-in-technical-alignment-is There’s also this: https://www.lesswrong.com/tag/ai-alignment-intro-materials
If you ask a typical alignment researcher about HCI, they will most likely say that the user interface is one of many places where an advanced AI system could fail critically. Typically researchers are more worried about critical failures elsewhere, since advanced AI systems present special challenges. Imagine a powerful AI system, no one is quite sure how it works. It turns out that somewhere in its massive thicket of cognition, there’s a subcomponent that can be characterized as having a goal, that goal isn’t compatible with human goals, and the AI system is smart enough to manipulate humans in order to achieve its goal. It’s not clear to me how better user interface design could address this. It seems like mostly a technical ML challenge. Nonetheless, I welcome HCI people learning about alignment and thinking about how ideas from HCI could help. Even if they’re mostly working on the user interface, such work could be really valuable.
Broader than HCI? Is that even possible? From a certain perspective HCI has subsumed much of the humanities as this point. Suggesting that HCI is just about the final interface layer would undoubtedly offend many people in HCI, maybe even more than the last part of my post already has. :-)
I get that alignment problems as approached in ML are interpreted very differently (often more technically) than much of what HCI does. As I imply in the post, I don’t think HCI is going to solve many of the specific problems ML has lumped under alignment, just that at a high level, core HCI sensibilities, such as the conviction that the values and needs of potential users need to be integrated at every possible step, are about trying to achieve alignment throughout the design and development process.
I don’t really find it surprising that HCI and ML alignment share a similar-seeming high level goal but go about trying to achieve it very differently. This seems inevitable when the goal is so broad.
I suppose if someone wanted to motivate HCI for ML people, it might be helpful to provide some compelling case studies from HCI where looking at a project through an HCI lens required throwing away a lot of work and redoing the project from the ground up. Perhaps there’s already a book or paper which collects such case studies?
There are probably many with the flavor that if you don’t try to understand the tasks you want to support first, you can design something useless. In visualization research for instance there is a “nested model” for design and evaluation of vis systems which is meant to ensure that you get things right at each layer of decision-making, starting with not getting the problem or key task abstractions wrong. A specific example that comes to mind is building a vis system for global health experts dealing with outbreak data and thinking that if uncertainty matters, it would be visualizing quantified uncertainty but then realizing the bigger source of uncertainty is differences in data quality by location, which is only really accounted for in the experts’ head (https://ieeexplore.ieee.org/abstract/document/8449328).
But it seems a challenge with finding these cases is that often in HCi people are doing user-centered design to begin with, which makes it harder to misdiagnose the task or problem and less likely that you end up building something unsalvageable. Like the example i just gave was never really a failure of the system, more of the initial design intuitions, because the authors used a heavily user-centered approach.
I suspect there may be more relevant examples in venues like FAccT, maybe AIES, which have some resemblance to CHI in terms of multidisciplinarity but are kind of their own genre (with more AI/ML overlap). E.g., this paper from last year’s FAccT discusses examples where predictive optimization in cases like predicting risk for making criminal release decisions, deciding who to hire, welfare referrals etc. is argued to have been a bad idea to start with and can’t be fixed by designing a better interface: https://dl.acm.org/doi/pdf/10.1145/3636509
Hence my conclusion that if HCi people want to have impact on ML, they should invest more time in studying ML. I think this is a general point about collaboration actually: if you want your skills to have impact on some other field, you should probably assume that you will have to become an expert in that thing yourself. Even if you don’t end up making it that far, my experience is that that attitude is often necessary to collaborate with people that come from very different areas than you are trained in.
“..Unfortunately though, I don’t really see much incentive for the average ML researcher interested in alignment to invest in the HCI way of doing things…”
In a large corporation, especially in a regulated industry, there are growing incentives to have a robust framework for reviewing all sorts of potential harms from deploying ML. Some of these potential harms are very UX centric like communicating confidence/uncertainty in a prediction, other harms are the ‘usual’ foundational biases in the ML workflow and problem definitions.
One particularly thorny issue occurs when the product team insists that publishing training materials indicating that human review is necessary (please verify the result, do not solely trust the result of the prediction, …) but these caveats are easily ignored soon after exposure to the product. How does one design a system where the humans actually need to or at least are encouraged to perform these reviews?
>How does one design a system where the humans actually need to or at least are encouraged to perform these reviews?
This seems like it would be hard to achieve in general – if there’s no incentive to verify in whatever setting the person is using the model for, then how can providers of these models expect to change that?
I just went to an HCI conference where I saw a few front-end tools designed to surface randomness in outputs to a prompt. Seeing variations on the response won’t really make much difference if the same misinformation appears in all of them, but probably a good default.