Treating AI review like the contentious policy design problem it is

This is Jessica. Many researchers are thinking about what we should do about scientific peer review now that AI makes producing papers so much easier. Submission numbers keep getting higher — in the past week, I saw reports that the most recent ACL submission cycle got 17k+ submissions, up from ~10k last cycle. TMLR went from getting 500 submissions every 60 days or so to getting the same number ever 19 days. There are simply not enough human reviewers to handle the surge, at least not without a dip in quality. The noiser the review system gets, the greater the incentive to submit sloppy papers, because you might get lucky. This is the so called “review death spiral.” 

It is a hard problem. Quotas on submissions per author are one avenue forward, which TMLR just announced it would adopt. Not surprisingly, many reviewers are also turning to AI to help. The question becomes how to design AI review protocols to help reduce some of the noise, through preliminary filtering or flagging or helping guide human attention to parts of a paper that are most likely to be problematic. 

But what sorts of checks should an AI review assistant run on a paper? It’s useful to separate basic integrity violations AI could flag, like is there evidence of plagiarism, fake citations, missing code/data to reproduce main results (which are comparatively less controversial) from “epistemic filters,” like does the paper pass replicability checks, robustness checks, preregistration checks, statistical significance checks, etc. There’s a temptation to blur these things in proposing how to apply AI to review. It’s easy to assume that the metascientists have already established that practices like replicability or preregistration are truth-indicating and we can just implement them at scale (and indeed, ML researchers are citing open science and other reform arguments to back their proposals).

But if there’s one lesson to be learned from the aftermath of the replication crisis, it’s that there is no small, stable, non-conflicting set of detectable signals of good science that will find the good stuff and reject the bad. There are heuristics that can be useful prompts for deliberation – get in the habit of preregistering, make sure you can replicate your results, test the sensitivity of your results to choices you made along the way – but things get weird when we start treating them like universal requirements. Authors shift attention away from unrewarded signals, like better theory or exploratory work, and become preoccupied with rigor signaling through their methods. The result is not necessarily more thoughtfulness. 

And so even if the AI review tools we create are simply intended to inform human reviewers about what checks a paper passed, what we implement will have important policy implications by incentivizing more work like that in the future. I don’t think we are in a good position to predict what happens if suddenly we require multiverse robustness or statistical significance in a field like machine learning, which has in many ways been all about iterative improvement and “frictionless reproducibility” rather than individual results passing all the robustness checks.

The answer is not to avoid using AI in review until we can find a non-gameable set of credibility qualities to have AI focus on, as some have recently argued (though I agree with the linked paper that we need more rigor in how we go about motivating review tools). Non-gameability sounds nice, but any automated review policy that allocates attention will be gameable, because ensuring good science is not so simple as finding the right checklist. The relevant question is instead what assumptions and downstream incentives we are willing to tolerate. To this end, at the very least we should get in the habit of spelling out the assumptions we’re making, so that the trade-offs of focusing on particular proxies become explicit.

I wrote up this view recently in a paper called “Stop Treating Metascientific Heuristics as Quality Filters in AI Review.” Here’s the abstract: 

AI-implemented checks for reproducibility, robustness, preregistration, claim scope, and other intended proxies for scientific credibility can extend human reviewers’ capabilities. However, treating metascientific heuristics–whose theoretical grounding remains contested or incomplete–as necessary and sufficient signals for filtering out bad science is counterproductive to scientific progress. The emerging literature blurs the line between integrity filtering, based on necessary but insufficient signals of validity like reproducibility of stated results or lack of fake citations, and epistemic filtering, which uses machine-detectable signals to judge scientific quality. Drawing on critical metascience, we show that commonly proposed signals of research quality are insufficiently justified as general indicators of scientific value. The answer is not necessarily to ban AI in review, given the deluge of submissions venues are facing. Instead, in recognition of how any use of automated signals–even when deployed with human oversight–will shape attention and create incentives upstream, developers of AI review tools should explicitly specify their assumptions about how proxy signals inform on scientific quality in the context of specific review decisions. This approach treats AI review contributions as contestable decision policies that will shape future research, acknowledging the value-laden nature of scientific judgment and surfacing relevant tradeoffs. 

Rather than arguing for or against any particular proxies, I’m more interested in the methodological and philosophical mindset we should bring to the new questions raised by AI review. To demonstrate what I mean by more explicit motivation, I analyze an example review decision problem and set of detectable signals in the appendix, drawing on an analysis of how statistical significance and exact replication success relate to signal-to-noise ratios measured under error from a recent paper by Eric van Zwet, Andrew, and Witold Więcek. The takeaway is that the value of a proxy will depend on how you define the latent state you care about (e.g., whether the direction of an effect was correctly estimated, how big the true signal-to-noise ratio is), what you assume about the generating process (i.e., how the proxy noisily reflects the latent state), and what you assume about the decision-maker’s choice of actions and utility function. By suggesting this approach, I am *not* suggesting that one can validate a new review tool’s utility before its been deployed. The point is that there will be trade-offs no matter what, and the best we can do is be concrete about the kinds of  assumptions that have to hold for proxies to be useful in review, so the community can debate what risks they are willing to accept. 

In this sense, my argument is very much along the same lines as Devezer et al’s argument that those proposing reform procedures should adopt more formal methodology to avoid unwarranted overgeneralization. Once checks become part of review infrastructure, they stop being neutral diagnostics and become policy levers. Let’s start treating them as such in research on AI review.

When is detecting AI-generated text worthwhile?

This is Jessica. AI-text detectors are coming to play a bigger role in adjudicating what texts are worthy of our attention. There was the surprising case of an apparently AI-generated short story winning the Commonwealth Foundation Short Story Prize, which returns 100% AI generated by Pangram, the leading detector whose false positive rate is reported as roughly 1 in 10,000 in its own audits and near zero on medium-to-long passages in an external audit. Applying Pangram to the other 4 stories that won awards this year suggests two others were heavily AI-assisted. More recently, the NeurIPS Position Paper track announced that it was desk rejecting 18% of submitted papers that were detected by Pangram as fully AI-generated. Another 13% are getting followed up on with the authors to investigate AI use. In this case the Call for Papers made clear that submissions should be “substantially written by human authors,” so this should not have come as a surprise.

We’re having to reconsider what authorship means. Can a person create literature or express their position on a subject without writing a single sentence themselves? When do we really care who strung the words together?   

Some people think detection is a waste of our collective time because we will never reach an equilibrium. AI-generated text will keep shifting toward what passes the detector. Human writers will continually update their beliefs about what features are indicative of AI-writing, but will also be influenced to write more like AI by reading so much AI text. There’s no stable target, just an endless cat and mouse game that incentivizes being savvy enough at any given time to avoid getting flagged. Meanwhile people are being morally scorned and suffering reputational damage for being caught on the wrong side of things. This may disproportionately affect some writers (like non-native english speakers) who are finally seeing the playing field leveled a bit. 

On the other hand, there are situations where it really is important to know who strung the words together. Education is the most obvious one. It’s just very hard to teach someone to think if they’re not writing down their ideas themselves. 

The problem is that outside of select scenarios like teaching, what we really tend to care about is who controlled the ideas, and this is not equivalent to who strung the words together. Some would argue that the latter is becoming increasingly irrelevant given that AI can write more fluently than many people and many people prefer AI-generated text. 

Of course the reason we’re seeing detection used to filter paper submissions is because the ideal process–where the content of each paper is carefully considered on its own merits–is increasingly untenable given the huge surge in submissions in some fields. It’s easy to pump out credible-seeming papers with minimal human oversight using AI, and enough people are doing this to create serious problems. 

Mostly my response is that if we are going to debate the value of detection we should be willing to make our assumptions explicit. So let’s walk through a toy model to think about what we’re really conjecturing about.

One way to think of the latent state that we actually care about in paper review is the author type. Let’s say type A authors come up with their ideas and do a lot of the writing themselves. Type B authors rely on AI to do much of the thinking for them, and also use AI to do much of the writing. Type C authors come up with their own ideas, but engage in extensive prompting to get AI to write everything they want to say for them.*

For each paper, we choose to either pass or reject, conditional on the output of a Pangram check. Let’s say we only care about whether it flags 100% AI generated or not, so the signal s is binary, where s=1 means AI detected.

Based on available Pangram audits, if a text is actually written heavily by AI there is a very high chance it flags as AI-generated: beta=P(s=1|AI written) with beta very close to 1. If a text is not written by AI, there is a very small chance it flags as AI-generated: alpha=P(s=1|human written). Pangram’s internal audits put alpha around 10^−4 but other audits find essentially zero false positives for medium-to-long passages. 

So P(s=1| A)=alpha, and if we assume Types B and C use AI to a similar extent for the writing, then \beta=P(s=1|B) = P(s=1|C). The posterior probability that a flagged paper is from a Type B author is then:

P(B|s=1) = (beta × p_B)/(alpha × p_A + beta × p_B + beta × p_C), and since alpha is tiny and beta is close to 1, P(B|s=1) ≈ p_B/(p_B + p_C)

The relevant considerations become what we think the author population looks like, and how costly we think a false positive versus a false negative are. 

As a starting point, let’s say that for our conference submissions this year, Type C is the rarest, at 20%, and Type A and Type B equally split the remaining mass at 40% each. Let’s also say that we consider rejecting an acceptable paper, c_FP, to be twice as bad as passing an unacceptable one c_FN. 

The optimal decision rule is to reject if c_FN​ * P(B|s=1)>c_FP * ​P(A or C|s=1), or equivalently P(B|s=1)>c_FP/(c_FN+​c_FP​​)

With c_FP=2 and c_FN=1, this means we reject if P(B|s=1) > 2/3.

Under the prevalence assumptions above, P(B|s=1) is approximately 2/3, so we are right on the boundary. From the standpoint of making the right decisions for this particular conference cycle, it’s not obviously bad. But if Type C is a little more common, e.g., we shift a little mass from p_A to p_C to make p_C 0.25, then P(B|s=1) is 0.62, then we shouldn’t desk reject only based on the flag. Similarly if we were to decide that falsely rejecting an acceptable paper is three times as bad as passing an unacceptable one, we shouldn’t rely on it alone. 

This model is obviously very simple. But it shows us what kinds of things we have to make assumptions about in the most basic case. Obviously I don’t really know how many people are using AI blindly to write papers, nor how many people are relying heavily on AI to write up their own ideas. You should take my numbers with a grain of salt. Personally I can’t imagine how relying on AI to do all the writing when I came up with the ideas would ever feel efficient, because I tend to have strong opinions on how things are said. But I can accept I am probably more of a control freak than many others. And AI overreliance is easy to slip into. Maybe papers chairs from recent ML conferences (or arXiv moderators) have estimates on bad-actor rates based on what they are seeing. 

What this exercise can’t tell us is how scientific progress is impacted by the warping of incentives that can happen when we use AI-detection as a filter. Classic principal-agent problems suggest that when we care about something hard to observe—like scientific quality or long-term epistemic value—but must rely on observable proxy signals to judge authors’ outputs, we should expect authors to shift more effort toward improving exclusively on those proxies. Avoiding m-dashes and ‘not this, but this’ constructions and whatever else currently ups the posterior probability of AI-generation is orthogonal to the actual thinking that research requires. What if relying more heavily on AI to write up our ideas is a good idea for science in the long run, in terms of more clearly communicating the ideas or saving a lot of time, so that we can get more good ideas out in the same amount of time? Then too much emphasis on detection might slow us down. However, I’m doubtful we are currently anywhere near a state of the world where discouraging writing with AI is as costly for scientific progress as spending time reviewing and reading many more questionable AI-generated papers is. The bigger threat at the moment is the slop overwhelming our ability to find the good stuff.

*We could also posit Type D authors that get AI to generate the ideas, but then write the papers themselves to evade detection, or are extremely good at getting AI-written text to evade detection. But this seems much less likely so I’m ignoring it.

What if scientists really were dispassionate observers, communicating ideas without irrational commitment? Look here, says AI.

This is Jessica. We often idealize science as proceeding primarily by the scientific method, where scientists approach the objects of their investigation with a healthy dose of detachment and neutrality, who become convinced only when the evidence is there, and remain open to changing their mind if new evidence becomes available. But in reality we see examples of authors becoming personally attached to their ideas despite the data, slipping into advocacy and becoming defensive or going into denial mode when presented with clear evidence they were wrong.

The seemingly irrational attachment to the ideas or findings can seem easy to dismiss as a bad thing. Yet there are also times when having some level of personal commitment makes one more effective at certain roles that scientists must play. For example, being too transparent about our own uncertainty is not always effective when presenting research to others, because the audience can become distracted and stop listening entirely, even if you have some useful insight to convey. My question is, What does our ability to now use AI to generate implementations, presentations, and even the ideas we work on themselves add to the mix?

I got a glimpse of this recently. May ended up being workshop month for me, with at least one each week. I saw a lot of presentations. A couple of these showed me something I hadn’t yet seen, at least outside of student presentations: talks comprised of obviously AI-generated slides. If you’ve tried to use the non-design optimized versions of models like GPT or Claude to create slides yourself, you will know what I mean. Almost every slide has content organized in a grid. There’s too much text—full sentences or nearly so in multiple places, headers and footers, and stylized phrasing everywhere, like “principal design levers” and “load-bearing assumptions” and “actionable pathways”. 

These were not presentations by overwhelmed junior faculty or researchers I’d never heard of. They were by prominent researchers who are respected in their fields. 

Needless to say, they were not very effective talks. The slides tended to have too much going on to parse in time, with way too much text. The vague phrasing was distracting, making me wonder what exactly the presenter meant by terms like “governing frictions” or “strategic bottlenecks” and whether they write like that in their papers too. Part of the problem is that the presenter tends to use their own language as they present, rather than reinforcing what’s on the slide, so you have two competing streams of information that feel like they’re from two distinct viewpoints, one which is quite confident and willing to summarize and even exaggerate, the other more reserved. 

It makes sense that you’re more likely to hold at a distance what you didn’t come up with yourself, subconsciously at least, even if you think you’re selling it. In one case, the speaker also described how some of the results themselves were discovered by AI, which probably further contributes to the impression that they hadn’t fully committed to what they are presenting. 

This has me wondering what the impact on diffusion of ideas will be as it becomes more standard practice to rely on AI for implementation in scientific production and communication. It’s funny how reserving skepticism for your own results often comes up in discussing epistemic virtues, but when speakers present as if holding their work at arm’s length, the result is not so informative. As we rely more heavily on AI in all stages of research, will we face more challenges in getting others to adopt our ideas? 

It’s also another reminder of how few people thinking about AI for science seem to have considered all the personal stuff that goes into the practice of science, with lots of irrational investment and fixation and stubbornness and pride to drive the loop of discovery and validation and communication. Scientific discovery may be an “ocean,” to borrow an analogy associated with Leibniz, but surfing it requires strapping oneself to a board and committing to seeing where it gets you, not just keeping it in sight while you splash around somewhere else. 

This also leads to a practical question of how you instill a sense of ownership, or at least commitment, to ideas that were partly produced by AI. My own experience is that it takes a lot of time to verify AI produced results before I get to the level of confidence I’d have if I’d done it myself. For complex tasks there will inevitably be decisions made along the way, e.g., about how to parameterize certain things in implementation or to deal with edge cases or other exceptions. Each of these has to be reconstructed before I can really feel that I stand behind the output. 

Is there an alternative? It makes me think of the “baking guilt” that housewives supposedly felt after cake mixes came on the market, because they only required adding water. There was a loss of a sense of personal contribution and emotional ownership. The solution, which persists today, was to have them add an egg. Some psychoanalysts went so far as to interpret this as symbolic of their fertility. For AI-aided science, the closest thing to adding an egg seems to be having agents explain at length to you what was done, which can still mean a big improvement over implementing everything yourself, but not as much of a boost as it first seems.

At any rate, interpreting the new challenges of AI-generated presentations of potentially AI-generated ideas as an aesthetic problem, or of “putting style before substance,” does not seem right. Scientific ideas don’t diffuse as bare propositions. They diffuse through people who have developed some passion for them. If we’re talking about AI for science, we shouldn’t be ignoring scientists and their relationships with what they do.  

The “humans are imperfect reporters too” defense for ascribing little thoughts to machines

This is Jessica. In my last post about the tension between the necessity that we ascribe human folk psychological concepts like thinking and reasoning to machines and the problems that arise when we overinterpret them, I briefly mentioned a defense I sometimes hear for anthropomorphic terms. The defense goes like this: We can’t protest the application of words like reasoning or belief or desire or understanding to AI because humans are also not always trustworthy reporters of the latent states that these terms refer to. And if humans can’t demonstrate full awareness of their reasoning process or their beliefs or their understanding or their intentions, why should we expect AI to? So, the argument goes, it is asymmetric to require complete faithfulness of LLM traces to some latent state. That would be setting a higher bar than we currently use with humans.

This line of reasoning takes two things that may resemble each other on the surface (e.g., what humans report their reasoning or belief or intention to be, and what machines report) which we have no solid reason to believe are produced by similar processes, and says they cannot be distinguished because we expect both to be reported with loss.

It’s bad logic. But more than that, it occurs to me that by doing this we indirectly deny the value of our subjective, non-verbal experience. It’s true we may not be able to faithfully report how we reasoned or what exactly we desired or intended or understood. But part of the reason we have terms for these things is because we perceive a distinction between more genuine and less genuine versions of them within ourselves. These qualitative distinctions are how we come to believe that reasoning and belief and intention and desire exist in the first place: because we can perceive them as being authentically present to different degrees in different situations, even if sometimes we seem to be mimicking the real thing rather than really doing it. It’s what licenses researchers to keep trying to study these things in people despite the difficulty of getting an unbiased read. I’m not saying we shouldn’t look for analogous processes in machines. But we should acknowledge that the referent for these terms is grounded in subjective experience. Among humans that has worked fine, because we assume ourselves to share an understanding of what it’s like to have an inner life.

Andrew recently analogized chatbots with being in autopilot in conversations that occur in a meeting, where if you have a rich enough memory bank of observations related to what the speaker is saying, you can spout appropriate conjectures and feedback without having to exert much conscious effort. But I think fake thinking extends to much more than just being on autopilot in conversation. It can seep into all of our decisions, whether about research ideas or methods or life decisions like what job to have or where to live. It’s very tempting to live by patterns, even when it’s not what you feel you actually want.

Our perceptions that there are more real and more fake versions of our own thinking or reasoning or belief or desire is what makes us feel more “alive” or “awake” in some situations over others. It connects us to the present. Only when we try harder to do these things authentically or recognize the motives that are actually driving us (despite the explanations we might want to assume) do we feel like we are really living, rather than just performing some role. And so our ability to perceive the difference seems intimately connected to the process of understanding ourselves.

From this angle, it is unfortunate how much of the AI rhetoric we’ve come to take for granted (at least in machine learning) — i.e., AI as scientist, AI as decision-maker, etc.— enacts this implicit move of equating human and machine processes by their outputs, and consequently subtly devaluing the role of our internal reality in giving these terms meaning. We take examples from the few domains where language can fully capture reasoning–math, coding–and we reduce all reasoning or thinking or intention to what can be made manifest.

Maybe this helps explain why there’s so much emphasis lately in certain circles (including tech) on being “high agency.” We become more insistent about shaping our external world as we lose appreciation for our internal one.

“As our daily lives involve ever more sophisticated computers, we will find that ascribing little thoughts to machines will be increasingly useful in understanding how to get the most good out of them” but “we must be careful not to ascribe properties to a machine that the particular machine doesn’t have”

This is Jessica. Maybe one of the biggest crimes of academic computer science (besides routinely ignoring prior work and making up social science to suit our needs) is our tolerance for abuse of language. We take technical things and inject them with social significance without thinking through what we’ve implied. This is perhaps forgivable in early stages of research when we’re trying to get more people excited about exploring some direction, but at some point people start taking things more seriously and we find ourselves committed to terminology that overreaches. Then the question becomes what, if anything, we should do about it.

Previously it didn’t feel like such a crime to talk about intelligence or learning in machines because nothing really worked that well, so the labels were clearly aspirational. But now it’s much easier to believe the simulacra. And so it becomes harder to tell when we are using human-oriented terms as a predictive convenience versus a scientific claim versus a marketing device. There are ramifications of referring to models’ reasoning or beliefs or chain of thought or explanations or intentions. Lots of people—from end users having personal relationships with models to media and AI companies themselves referring to “parenting” the latest models or asking if they can be “children of god”—are taking models too seriously. It’s bad enough in a computer science context that I now take for granted that if I want to refer to participants or scientists or decision-makers, unless I mean AI, I should add “human” in front, because otherwise the audience will assume I mean AI agents. Someone reminded me at a workshop recently how silly all this sounds to people who aren’t used to it. 

Too much casualness with words is unscientific. There was no good reason in the first place to call the token sequences a model produces when we ask it to “explain its reasoning” reasoning, other than that’s what we wish we could see. What an LLM is doing is distant from what happens when a human thinks about something, even after all the RL post-training. Similarly, we call lots of things “explanations” when we have barely begun to figure out what causal evidence we’d need to see to claim the output faithfully explains the model’s process of arriving at it.

But it can also seem unscientific to simply declare that “only humans can have beliefs” or “reason” or “provide rationales.” There’s no non-arbitrary line that we can draw between systems whose makeup or behavior truly warrants applying constructs like beliefs and desires and those where it is simply convenient to act as if they have these qualities. If you’ve ever tried to take beliefs seriously, from a decision theoretic perspective, you quickly come to realize that the “real” beliefs we assume a person has are a mythical thing that we will never directly observe, and you fall back on equating “beliefs” with a simpler idea: the probability distribution we arrive at after an elicitation process.

Much has been said in defense of a functional perspective toward using folk psychology terms with machines, where we decide what’s appropriate based on the predictive validity of the terms for our own understanding and use. John McCarthy wrote in 1983 that anthropomorphism can be a good idea “when it says something that cannot as conveniently be said some other way.” He argued that ascribing mental qualities and processes to machines helps us “understand what they will do, how our actions will affect them, how to compare them with ourselves and how to design them.” Perhaps the best reason to do this is that we get to draw on our existing familiarity with what phrases like “wants” do and do not convey; e.g., we all understand that if we say “The dog wants to go out”, that doesn’t mean that the dog believes itself to be capable of wanting or even that it’s conscious of what it wants, there’s just a sense in which it is trying to get to the state of being outside. 

The functional perspective has led to some strong statements suggesting that it is not only valuable to apply psychological terms to AI, it is necessary. These arguments often refer back to Daniel Dennett’s distinction between three stances we can take to machines: the physical stance, which is about its physical levels of organization, the design stance, which involves understanding it in terms of the purpose it was designed for, and the intentional stance, where we try to understand it by ascribing to it beliefs, goals, intentions, likes and dislikes, and other mental qualities. For example, philosopher Keith Frankish argues that “it remains true that adopting the intentional stance is the only way of interacting with an LLM in any interesting way; indeed, an LLM-powered chatbot that couldn’t be viewed as an intentional system would be completely useless.” McCarthy writes that “Long before we can make machines with human capability, we will have many machines that cannot be understood except in mental terms.” Similarly, about ascribing to a machine the concept of “trying” to do something, he says  “If the machine may do something we don’t know about but that can later be explained in relation to a goal, we have no choice but to use `is trying’ or some synonym to explain the behavior.” 

But forty-plus years later, McCarthy’s piece reads mostly like a defense of a very basic kind of debugging value of trying to imagine a program’s “state of mind” in situations where there’s little risk that we’re going to start mistaking the program for human-like in other ways. He wants to claim that when a thing is designed to act as if it had a certain belief, it can be better understood and manipulated by assuming it’s capable of that kind of belief. But surely he would agree that if a person who is suicidal is interacting with a language model that speaks as though it fully understands the complexity of their situation and what is best for them, it still isn’t always in the person’s best interest for them to take for granted that it does understand.

The dilemma is how to lean into the intentional stance when it helps, but to avoid overreaching. This seems hard. When you first start programming, you realize how easy it is to assume a program is smarter than it is. We are not very good at recognizing when we have slipped from reasoning “as if” to projecting. When technology slots into our human vulnerabilities, like our fear of intimacy and desire for companionship “without the friendship“, as Sherry Turkle said, we are in trouble. Even McCarthy calls it problematic to assign emotional qualities to machines, at least in his time, because “We have enough trouble figuring out our duties to our fellow humans and to animals without creating a bunch of robots with qualities that would allow anyone to feel sorry for them or would allow them to feel sorry for themselves.” 

Another reason to think carefully about the language we use is that it may shape what we can imagine in the future. This last part of McCarthy’s statement–“without creating a bunch of robots that would allow anyone to feel sorry for them”–hints at how what we project onto machines can shape how we go on to create them. It certainly seems possible that aspirational labeling played a role in getting us to a point where we have models producing sufficiently human-like outputs to have us gushing about their thinking process. I’m reminded of the color perception studies by Berlin and Kay, who found that what color chips different populations could differentiate was predictable from what color terms were available in the vocabulary, as if what we can name defines what we can see. At one extreme, Lucy Suchman argues against unquestioning acceptance that “AI” itself is a coherent thing, because it reifies it as a category for future investment. 

For Frankish, the line should be drawn at assigning communicative desires to models; they are playing a communication game (human-like chat) for the non-communicative reason that they are trained to play that game. Ascribing this single desire is enough to get us all the predictive power of the intentional stance. Consequently, we make a category error when we fall prey to stunts like giving a retired model its own blog because it requested it: we should expect an LLM to say things like this because it is designed to roleplay. Elsewhere I call this kind of mutual sympathetic relationship with AI “idiot compassion,” a phrase borrowed from Buddhist monk Chogyam Trungpa.  

But couldn’t humans also just be playing a chat game? Why is it ok to say a human is reasoning or has intentions or desires, when we don’t know exactly how those concepts map to observable physical processes?  Frankish argues that our linguistic behavior is corroborated by various non-linguistic sources in a way that LLMs’ is not.  He talks about a difference in our ability to hold epistemic stances toward statements (what Dennett would call “opinions”) from that of LLMs. We can conceive of assigning different levels of credence to statements. I can repeat something someone said verbatim without believing it, or I can be fully committed to the truth of something I say in the sense that it will guide my future behavior. LLMs can do this too, but their opinions lack grounding “in a web of non-linguistic behavior in which a wider range of desires can be attributed.” It’s a shallower form of epistemic stance, at least at the current moment of development.

I like AI, but I don’t like contributing to thoughtlessness. Better semantic hygiene seems warranted, even if it seems like the ship has already sailed. We could shift emphasis to the interpreter (us) by referring to the “human story” or “human pleaser” or “anthropomorphism fulfiller” instead of the chain-of-thought or reasoning or thinking trace. Or we could just add “fake” before whatever humanization we prefer, i.e. the “fake thinking trace,” or “fake reasoning.” I also like “so-called reasoning,” like I like “so-called replication crisis” as a way of pointing to a concept while questioning the expectation. 

P.S. Thanks to Manesh Agrawala for a conversation that inspired this post. 

P.P.S. Some of these McCarthy quotes appear in Recursion, the play Andrew and I wrote!

Show me science

This is Jessica. Lately I’m thinking about how AI review changes the scientific evaluation process, and by extension what authors are incentivized to report. Some speculate that the scientific paper, as a summary of the research for other humans, may be on the verge of becoming obsolete or at least less important. Instead, we may see more raw summaries of just the facts. The idea is that if LLMs are increasingly the consumers of research, we don’t really need all the baggage of narrative and illustration to make things more relatable to humans. 

Last November, Tom Dietterich asked for opinions on social media about what arXiv should do about papers that are bulleted lists, like this:

bulleted list of results with little narrative

Some of the impetus to reduce narrative in papers predates LLMs.  E.g., even before the appearance of chatbots in 2022, people were arguing that we should get rid of or restructure Discussion sections in scientific papers, because authors are often tempted to use them to drift into rhetoric (“spin”) and unwarranted speculation. But most papers still include them. Maybe LLMs will be the impetus that actually shifts the norm.  

However, there are lots of ways to contextualize scientific contributions that are not simply rhetoric, and which are already disincentivized more than they should be. Some move in the opposite direction from adding interpretation, such that omitting them is like withholding the information readers need to judge the work. 

For example, one of my pet peeves with the way many AI and machine learning papers get written is that showing examples of the task is de-prioritized in order to fit in more results, especially in the main paper text. It’s very hard to evaluate how much improvements in performance matter if you aren’t given a single concrete example of the problem being solved! You made an LLM reviewer that’s great at finding errors in papers? Show me some examples of the kind of errors it’s detecting, so I can judge how much this moves forward our ability to verify science. You made a benchmark for the fairness of visual language models? Show me what the image and text prompts you’re testing the model on look like, so I can judge whether I agree that you are evaluating something meaningful versus a few people’s conception of what is politically correct. Instead we get high-level verbal descriptions of the kind of task (“detecting errors”, “fairness”, “content moderation”) and/or references to datasets, followed by lists of metrics and comparisons of how different models or algorithms performed. 

What’s weird is how comfortable entire fields can become going through the motions of evaluation without treating the task itself as part of the science. But authors are incentivized to bury the details by tight space limits on the main text, and to avoid giving reviewers more to pick apart. 

For us as human readers, seeing the examples often prompts a kind of common sense judgment about how “real” the problem is, making it harder for authors to pass off research that makes progress on made-up problems. But how much a task is likely to matter in the world is not the kind of thing current AI models are particularly good at evaluating. This makes me think of lots of other “just show me…” guidelines that tend to improve people’s ability to assess science:

Show me the plot: Don’t give me a big table of numbers, plot your coefficients so effect size and uncertainty dominate over significance (though you can still easily get that if you want by seeing which estimates include 0).  

Show me the variance: Don’t just show me the uncertainty in parameter estimates, plot the measurement variation. We found this reduced overestimation of treatment effects by lay people by a fair amount in this paper, and my coauthors also found similar effects with experts.   

Show me the interface: When gathering data from humans (whether ground truth labels for training or aligning a model, or behavioral responses to some experimental task), show me what they saw and how they were asked the questions, so I can judge how hard their task was, what biases might arise, etc. 

Show me the prompt: The LLM version of the above. If it’s long you may not have space in the main body of the paper, but all the prompts you used should appear somewhere.

Show me the failure cases: Seeing what instances throw a model or system off says a lot about how much progress has been made, and how the model may be succeeding in the other cases.

Show me the baselines: Improvements are meaningless if the reader has no idea what’s being improved over. Give the reader a bit of intuition about how the other approaches work, including the dumb ones you should definitely be beating. 

Show me the design analysis: I was on an open science panel last week where someone asked what they should look for as a reviewer and what they should report with their own experiments to help readers evaluate them, beyond open data and code. I said that something I often request in reviewing empirical papers is information on how the authors chose study design and sample size, including what effect estimates they were prioritizing with what intended level of precision or power. It’s easier to make sense of what’s been learned if you know what the authors were attempting.

Show me the raw output: Whenever you are doing qualitative coding of model outputs (or human responses for that matter) show me a couple examples of the original texts for each possible code. 

Some of these may be informative for LLM reviewers as well as humans, in the sense of helping them predict consensus human judgment on the paper, but others (like plots instead of tables of numbers) probably not so much. 

Fraud and the false optimism of AI for science

This is Jessica. “Scientific doomerism” seems to be everywhere lately, from a presidential statement that promises to restore “gold standard science” from the top down because scientists have botched things, to journals being inundated with AI-produced papers, to sleuths like Reese Richardson documenting the scale of organized scientific fraud through paper mills and collusion. In his post on this last example, Andrew wrote something that caught my attention: 

And, yes, typing some prompts into a chatbot and producing a paper is fraud, in the same way that publishing textbook excerpts as if it were new research is fraud, or copying from wikipedia as if it were new research is fraud, etc etc. It doesn’t require fake data and it doesn’t require some cackling Snidely Whiplash attitude. It can be some schlub sitting at a computer terminal who wants to get his contract extended or get admitted to a Ph.D. program or whose adviser is pressuring him to get some publications . . . But it’s fraudulent publication, not the same as bad research (which is actually research, it just happens to be useless because bad measurement and kangaroo).

There are clearly some differences between passing off work containing fake evidence you purchased from paper mills and work that you contracted AI to do for you–in one case you are paying money for someone to pass something fake off as real with your name on it, whereas in the other you might be using actual data and think you’re just saving time, especially if you’re reviewing what the AI does. So should we really consider both forms of fraud? 

It struck me how sharply Andrew’s perspective contrasts with the current direction of discussions among ML and other researchers interested in AI for science. There, it’s seen as inevitable and not necessarily morally problematic that the future of science will have humans largely playing the role of curators, who prompt and select among results produced by LLM agents, who do the bulk of the work. It’s worth considering what kinds of ethical lines this crosses exactly. 

Let’s imagine that I give an LLM an initial high level research question related to a topic on which I am knowledgeable. It churns on the idea and ultimately designs an experiment it’s happy with, I review the plan before prompting it to continue, maybe tweaking slightly, like changing a condition or suggesting an additional robustness check. It then gathers data on my behalf (e.g., running an online experiment or downloading existing datasets), conducts the analysis, and presents me with the results. I review these and then give it permission to write up a paper. I read the final paper to make sure I know what it’s saying before I submit. Maybe I change a few things I don’t agree with. I add my name and also credit the AI. In other words, there is a light human touch throughout, but much of what is presented as my work comes from the model.

From an “optimistic” AI-for-science perspective, the strongest argument is probably to cast it as part of the scientist’s job to try to make the most of current technology. If we think AI might help us be more productive, then we should explore how much time it can save us, just like it was a good move for statistics to embrace the computational revolution that made previously intractable models commonplace. Proponents of AI for science argue that it is irresponsible not to use AI given its current capabilities, just like it could be construed as irresponsible for a brilliant researcher to refuse to use calculators if doing so meant they could contribute more useful advances to the field. Of course, this assumes that we won’t be sacrificing anything vital in the process. 

The “pessimistic” view of AI generated science as fraud thinks we are sacrificing something vital in the process. But what is it exactly? If you believe that “the devil is in the details” (or “God is in every leaf of every tree”, depending whose side you want to be on) then whenever you outsource decisions you would otherwise make yourself, you have potentially compromised the work from the perspective of your own expert judgment. So putting your name on it betrays what you know to be true of good science. Of course, you could check everything down to the lowest level, and intervene whenever the agent tries to do something you don’t agree with. Then the AI is really just a means of computation–even if you use it for brainstorming what research questions to ask, you could view it as a way of extending your limited resources but without sacrificing your own scientific judgment. This requires that you are knowledgeable enough to assess everything it does. Assuming you are, then it seems hard to argue that this is fraud. Though admittedly, a lot rests on how careful you are when you check things over.

Part of the concern may be that AI makes it tempting to extend your methods or claims outside of what you know well. Without the option of using it, you would have had to do the research yourself (and presumably gain understanding in the process) in order to apply that method. Relatedly, I suspect most people would agree that when you are in a training context, like taking courses in grad school, turning in work that was largely driven by the AI is a form of fraud because it holds you back from gaining the understanding yourself. 

Another version of the fraud argument focuses on misattribution of ideas. Maybe this is what Andrew had in mind when he wrote “publishing textbook excerpts as if it were new research is fraud, or copying from wikipedia as if it were new research is fraud.” If the AI produced significant parts of the contribution, like the specific hypothesis, the choice of methods, and the framing of the contributions, then those aren’t your novel creations and adding yourself as author is misattribution. But this is tricky because human scientists also recombine existing ideas, methods, and frames constantly, and we often call the resulting combinations original contributions. So the question isn’t whether ideas are derivative per se but when the degree of derivation crosses into infringement. We have copyright law for some creative domains, and we can try to formulate when AI outputs are permissible in light of this (e.g., Annie Liang has some recent work on this). But to say AI-assisted papers are fraudulent on such grounds, we’d need to work out what the scientific analog of substantial similarity is. This is hard because science explicitly values building on prior work. But we can agree on some things, like you shouldn’t do something that is too close to others’ work without citing them.

A final angle on why its fraud might be that it misportrays what science is more broadly. If you think the reasons to do science are fundamentally human–that as scientists we are concerned with producing understanding for ourselves just as much as we’re concerned about improving things in the world–then you could argue that for science to be meaningful we have to be the ones coming up with the ideas and shaping them as we go. From this perspective, automated science isn’t inherently wrong, it’s just missing the point. AI for science arguments often completely overlook the “people production” role of science. In the extreme, they envision AI finding solutions for lots of real world problems and intervening to control outcomes in the world without us understanding how any of it works. In reality, the personal side–including the search for personal fulfillment through science–is a big part of why smart people who could potentially make a lot more money in applied roles end up choosing research careers. And it’s a big part of how we evaluate scientists. How many of your Ph.D. students have gone on to competitive research positions? What does the trajectory of topics you’ve worked on say about your research taste? 

Pushback to this argument might point out that by saying science is entirely a matter of human careers, we contradict claims that we as scientists like to make, about how we are dedicated to improving the state of the art in our field, or producing value for the world. Would it still be science as we know it if we started acknowledging that it’s really about personal fulfillment for scientists? But I think this is a bit of a false dichotomy. The public value of science depends on there being humans who find the work meaningful enough to do it well, including pushing back on sloppy results, exercising their taste to shape the direction in their field, training students worth training, etc. Careless AI use can threaten this by flooding the system with outputs that crowd out careful work, disincentivizing the people who would be intrinsically motivated to do quality work less likely to stick around. It also implicitly reframes science as nothing more than a pipeline for results.

My view is that AI use can go either way, depending on how you approach it. What best determines whether it’s fraud or not is the attitude you bring. It can help you do less fraudulent research if you’re the kind of person who is already very picky about what you send out to the world. But it can help you fool yourself and others if you let competitiveness and obsession with metrics drive how you use it.

As a final comment, there’s some irony in using terms like “optimism” to talk about this. I described the pro-automated science argument above as “optimistic,” because I think that’s how many in this camp see themselves–as fundamentally optimistic about the future of science and our ability to improve it by using AI. But the underlying motivation to figure out how to produce papers with as little human oversight as possible is also often deeply pessimistic. A common narrative is the “review death spiral”: AI production stresses the review system, which increases the noisiness of paper acceptance decisions, which further incentivizes submitting sloppy AI produced papers. The answer is presumed to be putting more AI in place on both sides. The idea that scientists have agency and could continue to shape the meaningfulness of what gets produced starts to seem out of the question.

Increasingly, a lot of the most enthusiastic pro-AI discourse (including for science) strikes me as nihilism masquerading as optimism. We have people who perceive themselves as huge optimists that will reshape science or society for the better simultaneously lacking the imagination to see beyond their own technological determinism. It reminds me a bit of the “optimism” associated with some open science and science reform positions, who also suggest that we just need the right technology to fix the problems (though in this case, its heuristics like replication or preregistration). It’s a fundamentally non-agentic view of human scientific endeavor. 

Epistemic Virtues for Science in the Age of Automation

This is Jessica. Back in the 1980s, novelist Italo Calvino developed a series of six lectures describing literary virtues he felt should be enduring regardless of how the world changed: lightness, quickness, exactitude, visibility, mulitiplicity, and consistency. These were published as “Six Memos for the Next Millennium” in 1993.

We aren’t on the verge of a new millennium, but recent advances in AI make automated evaluation and production of science increasingly possible. This raises the question of what qualities we should most be trying to preserve as processes change. It got me thinking it’s a good time to undertake a parallel exercise to Calvino’s, but for science. 

I enlisted Andrew and Berna, and as a first step we are seeking input on which “epistemic virtues” practicing scientists see as most critical to uphold. By virtues, we mean any durable qualities of scientific character and practice that shape how inquiry is conducted, claims are framed, evidence is evaluated, and disagreement is handled. At the link below, we compiled a set of candidate virtues for you to consider, and ultimately rank your top six. 

The virtues are Accountability, Apoliticalness, Authenticity, Awareness of Contextual Dependence, Coherence Seeking, Consensus Seeking, Curiosity, Discernment, Epistemic Cost-Benefit Awareness, Epistemic Fortitude, Epistemic Humility, Epistemic Pluralism, Impartiality, Indifference, Intellectual Autonomy, Intellectual Humility, Precision, Preference for Generality, Reproducibility Seeking, Reputational Grounding, Responsibility, Seeking Contestability, Seeking Correspondence to Observable Reality, Skepticism, Transparency, Unsettledness.

We provide short descriptions of each, and you can also tell us whether you think we missed any important ones.

We’re hoping for broad participation from researchers on this! If you are a faculty member, research scientist, postdoc, or senior Ph.D. student in any area of science, please take five minutes and fill it out. We’ll share the results widely along with some reflections. 

Survey link:

https://docs.google.com/forms/d/e/1FAIpQLSc_jHOrXpFMF3CKgHXfnUhZzeHLnagh3S1G5Kg8ZCyLPXUgxg/viewform?usp=sharing&ouid=103049774617868167713

Acknowledgments: Thanks to Carl Bergstrom, Pam Reinagel, and Tian Zheng for providing some of the candidate virtues.  

New course on generative AI for behavioral science

This is Jessica. It feels like an “old” course now that the quarter is almost over, but this winter at Northwestern I taught a grad seminar on Generative AI for Social Science. The goal was to survey emerging applications of generative AI (mostly language model agents) in the social sciences, with special attention to methodological and metascientific concerns that come up when AI is used to simulate or substitute for human observations or labels. I became interested in this topic last year as a result of the problems it presents to inference, but also the opportunities it may present to improve behavioral research, which we recently discussed here

I joined up my computer science section of the course with a Communications section led by my colleague Aaron Shaw, which resulted in a good mix of students across AI and social science. This was great for discussion, as many of the Comm students were experts in survey methods or psychology, while the CS students were more knowledgeable about how transformers work and methods for probing their internals. We also organized a workshop on validating generative AI for social science in February, which some of the students attended and therefore got to see the authors of the papers they were reading present them live.

The only downside of mixing backgrounds was occasional friction with the more formal methods papers. A few times in class there were dismissive jokes made about the more statistically-demanding papers, I assume as a result of people finding the content challenging. The course was advertised as requiring grad-level experience in stats, so they should have known what they were getting into, but I think sometimes people read these things and assume it doesn’t apply to them. My stance on these things is typically that as long as you’re doing your best to understand the material, it’s not a problem and I’ll help you get through it. But I don’t have a lot of patience with the attitude that because it’s challenging to you, it’s ok not to try to understand it. Especially in a class about why we need to take seriously the challenges that LLM simulations present for drawing valid inferences about human behavior!

This week they present project proposals, which include things like exploring new prompting architectures based in cognitive theory, using ML interpretability methods to steer models in ways informed by social science, and studying belief elicitation and uncertainty expression in language models. It’s a good topic for a seminar course because there are lots of aspects of methods that haven’t yet been explored, and thus lots of opportunity to bring social science theories to bear on how we interact with language models, or to apply the latest methods in AI or stats to behavioral science questions. 

Week 1: Course introduction:  Can generative AI transform social science?

This week we set context, reviewing proposals that argue for the transformative power of generative AI for social science.  

Optional:

  • Dillion, D., Tandon, N., Gu, Y., & Gray, K. (2023). Can AI language models replace human participants? Trends in Cognitive Sciences, 27(7), 597–600. https://doi.org/10.1016/j.tics.2023.04.008.
  • Anthis, J. R., Liu, R., Richardson, S. M., Kozlowski, A. C., Koch, B., Brynjolfsson, E., Evans, J., & Bernstein, M. S. (2025). Position: LLM social simulations are a promising research method. Forty-second International Conference on Machine Learning Position Paper Track. https://openreview.net/pdf?id=cRBg1dtj7o.

Week 2: LLMs as surrogates I: Attitudes, opinions, social behavior 

This week we start to read papers that evaluate how well LLMs can act as surrogates of humans. These readings focus on studies that use them to simulate human attitudes, opinions, and social behavior.

Optional:

  • Chuang, Y.-S., Goyal, A., Harlalka, N., Suresh, S., Hawkins, R., Yang, S., Shah, D., Hu, J., & Rogers, T. (2024). Simulating opinion dynamics with networks of LLM-based agents. In K. Duh, H. Gomez, & S. Bethard (Eds.), Findings of the Association for Computational Linguistics: NAACL 2024 (pp. 3326–3346). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-naacl.211
  • Hansen, A. L., Horton, J. J., Kazinnik, S., Puzzello, D., & Zarifhonarvar, A. (2024). Simulating the survey of professional forecasters (SSRN Scholarly Paper No. 5066286). Social Science Research Network. https://doi.org/10.2139/ssrn.5066286 
  • Park, J. S., Zou, C. Q., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Willer, R., Liang, P., & Bernstein, M. S. (2024). Generative agent simulations of 1,000 people (No. arXiv:2411.10109). arXiv. https://doi.org/10.48550/arXiv.2411.10109

Week 3: LLMs as surrogates II: Cognition and behavioral experiments  

This week’s material expands on the first “surrogates” set to focus on using LLMs to simulate human cognition and experimental effects more directly. 

  • Cui, Z., Li, N., & Zhou, H. (2025). A large-scale replication of scenario-based experiments in psychology and management using large language models. Nature Computational Science, 5(8), 627–634. https://doi.org/10.1038/s43588-025-00840-7.
  • Binz, M., Akata, E., Bethge, M., Brändle, F., Callaway, F., Coda-Forno, J., Dayan, P., Demircan, C., Eckstein, M. K., Éltető, N., Griffiths, T. L., Haridi, S., Jagadish, A. K., Ji-An, L., Kipnis, A., Kumar, S., Ludwig, T., Mathony, M., Mattar, M., … Schulz, E. (2025). A foundation model to predict and capture human cognition. Nature, 644(8078), 1002–1009. https://doi.org/10.1038/s41586-025-09215-4.
  • Tranchero, M., Brenninkmeijer, C.-F., Murugan, A., & Nagaraj, A. (2024). Theorizing with large language models (Working Paper No. 33033). National Bureau of Economic Research. https://doi.org/10.3386/w33033

Optional:

  • Chen, Y., Liu, T. X., Shan, Y., & Zhong, S. (2023). The emergence of economic rationality of GPT. Proceedings of the National Academy of Sciences, 120(51), e2316205120. https://doi.org/10.1073/pnas.2316205120
  • Ashokkumar, A., Hewitt, L., Ghezae, I., & Willer, R. (2025). Predicting results of social science experiments using large language models. Preprint.  https://docsend.com/view/ity6yf2dansesucf
  • Peng, T., Gui, G., Merlau, D. J., Fan, G. J., Sliman, M. B., Brucks, M., Johnson, E. J., Morwitz, V., Althenayyan, A., Bellezza, S., Donati, D., Fong, H., Friedman, E., Guevara, A., Hussein, M., Jerath, K., Kogut, B., Kumar, A., Lane, K., … Toubia, O. (2025). A mega-study of digital twins reveals strengths, weaknesses and opportunities for further improvement (No. arXiv:2509.19088). arXiv. https://doi.org/10.48550/arXiv.2509.19088
  • Akata, E., Schulz, L., Coda-Forno, J., Oh, S. J., Bethge, M., & Schulz, E. (2025). Playing repeated games with large language models. Nature Human Behaviour, 9(7), 1380–1390. https://doi.org/10.1038/s41562-025-02172-y

Week 4: Bias, alienness, and other threats to generalization

When the goal is to learn about human behavior, relying on LLM simulations risks biasing downstream inferences. The readings survey ways that LLMs tend to misrepresent human response distributions and exhibit non-human-like errors, as well as metascientific concerns that arise from their availability as a cheap source of simulated data.

  • Wang, A., Morgenstern, J., & Dickerson, J. P. (2025). Large language models that replace human participants can harmfully misportray and flatten identity groups. Nature Machine Intelligence, 7(3), 400–411. https://doi.org/10.1038/s42256-025-00986-z
  • Messeri, L., & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. Nature, 627(8002), 49–58. https://doi.org/10.1038/s41586-024-07146-0
  • Westwood, S. J. (2025). The potential existential threat of large language models to online survey research. Proceedings of the National Academy of Sciences, 122(47), e2518075122. https://doi.org/10.1073/pnas.2518075122
  • Mancoridis, M., Weeks, B., Vafa, K., & Mullainathan, S. (2025). Potemkin understanding in large language models (No. arXiv:2506.21521). arXiv. https://doi.org/10.48550/arXiv.2506.21521

Optional:

  • Atari, M., Xue, M. J., Park, P. S., Blasi, D. E., & Henrich, J. (2023). Which humans? (No. 5b26t_v1). PsyArXiv. https://doi.org/10.31234/osf.io/5b26t
  • Dominguez-Olmedo, R., Hardt, M., & Mendler-Dünner, C. (2024). Questioning the survey responses of large language models. Proceedings of the 38th International Conference on Neural Information Processing Systems, 37, 45850–45878. https://dl.acm.org/doi/10.5555/3737916.3739374
  • Wang, P., Zou, H., Yan, Z., Guo, F., Sun, T., Xiao, Z., & Zhang, B. (2024). Not yet: Large language models cannot replace human respondents for psychometric research (No. rwy9b_v1). OSF Preprints. https://doi.org/10.31219/osf.io/rwy9b
  • Cummins, J. (2025). The threat of analytic flexibility in using large language models to simulate human data: A call to attention. arXiv preprint arXiv: https://doi.org/10.48550/arXiv.2509.13397.

Week 5: Validation I

This week we turn our attention to the methods that authors propose to use to check how well a language model simulates human behavior and draws inferences about the world, or to get valid estimates of their predictive accuracy.

  • Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337-351. https://doi.org/10.1017/pan.2023.2.
  • Manning, B. S., & Horton, J. J. (2025). General social agents (No. arXiv:2508.17407). arXiv. https://doi.org/10.48550/arXiv.2508.17407
  • Vafa, K., Chang, P. G., Rambachan, A., & Mullainathan, S. (2025). What has a foundation model found? Using inductive bias to probe for world models. Proceedings of the Forty-Second International Conference on Machine Learning (ICML 2025), PMLR 267. https://openreview.net/pdf?id=i9npQatSev.

Optional: 

  • Neumann, T., De-Arteaga, M., & Fazelpour, S. (2025). Should you use LLMs to simulate opinions? Quality checks for early-stage deliberation (No. arXiv:2504.08954). arXiv. https://doi.org/10.48550/arXiv.2504.08954.
  • Larooij, M., & Törnberg, P. (2025). Do large language models solve the problems of agent-based modeling? A critical review of generative social simulations (No. arXiv:2504.03274). arXiv. https://doi.org/10.48550/arXiv.2504.03274
  • Aher, G. V., Arriaga, R. I., & Kalai, A. T. (2023). Using large language models to simulate multiple humans and replicate human subject studies. Proceedings of the 40th International Conference on Machine Learning (ICML 2023), PMLR 202, 337–371. https://proceedings.mlr.press/v202/aher23a.html

Week 6: Validation II

In contrast to heuristic approaches to validation that aim to show that LLM outputs are “close enough” to human ones, statistical approaches use some human observations to learn how to correct estimates drawn from LLM observations. These readings formally motivate why heuristic validation is not enough, introduce calibration frameworks, and present empirical results on how these methods compare to other repair strategies like fine-tuning.

  • Ludwig, J., Mullainathan, S., & Rambachan, A. (2025). Large language models: An applied econometric framework (Working Paper No. 33344). National Bureau of Economic Research. https://doi.org/10.3386/w33344
  • Broska, D., Howes, M., & Loon, A. van. (2025). The mixed subjects design: Treating large language models as potentially informative observations. Sociological Methods & Research, 54(1), 1074–1109. https://doi.org/10.1177/00491241251326865
  • Hullman, J., Broska, D., Sun, H., & Shaw, A. (2025). This human study did not involve human subjects: Validating LLMs as behavioral evidence. Preprint. PDF.  
  • Krsteski, S., Russo, G., Chang, S., West, R., & Gligorić, K. (2025). Valid survey simulations with limited human data: The roles of prompting, fine-tuning, and rectification. arXiv preprint arXiv:2510.11408. https://doi.org/10.48550/arXiv.2510.11408.

Optional:

Week 7: AI as social scientist

So far we’ve mostly talked about LLMs being used as plug-in simulations for human data In surveys and experiments. This week we broaden to consider use of AI in other parts of the research process, like identifying what to research or what to manipulate in an experiment. 

  • Manning, B. S., Zhu, K., & Horton, J. J. (2024). Automated social science: Language models as scientist and subjects (No. w32381). National Bureau of Economic Research. https://doi.org/10.3386/w32381.
  • Almaatouq, A., Griffiths, T. L., Suchow, J. W., Whiting, M. E., Evans, J., & Watts, D. J. (2024). Beyond playing 20 questions with nature: Integrative experiment design in the social and behavioral sciences. Behavioral and Brain Sciences, 47, e33. https://doi.org/10.1017/S0140525X22002874 (See also commentaries on this article).
  • Si, C., Yang, D., & Hashimoto, T. (2024). Can llms generate novel research ideas? A large-scale human study with 100+ nlp researchers. arXiv preprint arXiv:2409.04109.

Optional: 

  • Musslick, S., Bartlett, L. K., Chandramouli, S. H., Dubova, M., Gobet, F., Griffiths, T. L., … & Holmes, W. R. (2025). Automating the practice of science: Opportunities, challenges, and implications. Proceedings of the National Academy of Sciences, 122(5), e2401238121.
  • Tong, S., Mao, K., Huang, Z., Zhao, Y., & Peng, K. (2024). Automating psychological hypothesis generation with AI: When large language models meet causal graph. Humanities and Social Sciences Communications, 11(1), 896. https://doi.org/10.1057/s41599-024-03407-5

Week 8: Causal discovery & explanation

Continuing with the theme of using AI to design experiments or support theory, this week we look at using models or interpretability methods to learn representations of stimuli or outputs to support causal inference, discover outcomes in text, and theorize about human behavior.  

  • Imai, K., & Nakamura, K. (2025). Causal Representation Learning with Generative Artificial Intelligence: Application to Texts as Treatments. arXiv preprint arXiv:2410.00903
  • Modarressi, I., Spiess, J., & Venugopal, A. (2025). Causal inference on outcomes learned from text. arXiv preprint arXiv:2503.00725.
  • Zhu, J. Q., Xie, H., Arumugam, D., Wilson, R. C., & Griffiths, T. L. (2025). Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions. arXiv preprint arXiv:2505.11614. https://doi.org/10.48550/arXiv.2505.11614.

Optional:

  • Tak, A. N., Banayeeanzade, A., Bolourani, A., Kian, M., Jia, R., & Gratch, J. (2025). Mechanistic interpretability of emotion inference in large language models. Findings of the Association for Computational Linguistics: ACL 2025 (pp. 13090–13120). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-acl.679.
  • Movva, R., Peng, K., Garg, N., Kleinberg, J., & Pierson, E. (2025). Sparse Autoencoders for Hypothesis Generation. Proceedings of the 42nd International Conference on Machine Learning, PMLR 267:44997-45023. https://proceedings.mlr.press/v267/movva25a.html.
  • Zhu, J. Q., Peterson, J. C., Enke, B., & Griffiths, T. L. (2025). Capturing the complexity of human strategic decision-making with machine learning. Nature Human Behaviour, 1-7. https://doi.org/10.1038/s41562-025-02230-5.
  • Kim, J., Evans, J., & Schein, A. (2025). Linear representations of political perspective emerge in large language models. arXiv preprint arXiv:2503.02080. https://doi.org/10.48550/arXiv.2503.02080

Week 9: Belief-like representations and Bayesian inference

Behavioral scientists often take for granted that people have beliefs, attitudes, desires, and other mental states. This week we look at proposals for how to look for similar representations in language models. We also consider a Bayesian formulation of the prompting process that describes how the researcher’s expectations of reasonable data influence what they generate with language models.

  • Herrmann, D.A., Levinstein, B.A. Standards for Belief Representations in LLMs. Minds & Machines 35, 5 (2025). https://doi.org/10.1007/s11023-024-09709-6
  • Yamin, Khurram, Jingjing Tang, Santiago Cortes-Gomez, Amit Sharma, Eric Horvitz, Bryan Wilder.  (2026). Do LLMs Act Like Rational Agents? Measuring Belief Coherence in Probabilistic Decision Making. https://arxiv.org/abs/2602.06286.
  • Misra, S. (2025). Foundation Priors. arXiv preprint arXiv:2512.01107.

Optional:

Living the metascience dream (or nightmare) with AI for science

This is Jessica. We recently wrote about multiverse analysis, which takes the idea of sensitivity analysis to the extreme by multiplexing over every reasonable decision you could have made in analyzing your data. I pointed out that it’s now trivial to turn any paper into a multiverse analysis. Just ask Claude Code to create a notebook replicating the analysis in a paper, then add whatever variations you want. Or let it plan the multiverse too by suggesting some common ablations. Today I want to consider what this newfound ease in testing for robustness might mean for science.

There are a few corollaries of this new reality. One is that it’s now trivial to test reproducibility. No human reviewer may have the patience to deal with your messy codebase, but automated evaluation makes it much easier to detect the failures (at least where it’s possible for data to be made public) and prevent them from moving forward. This can take some curation of fallback plans for common failure cases (see, e.g., this recent paper), but once we endow agents with these skills, we can apply them at previously unknown scale. It’s a short hop from there to multiverse writ large.

Replicability testing at scale is also not far off. For example, you can have it synthesize data with the same structure but variation along some dimension. Before long we may also see agents being given permission to collect new data, for example by running online human subjects studies on Prolific or some other platform. The years of work it used to take to run a large-scale replication project may now translate to one grad student’s summer project.

In short, we should expect the level of scrutiny on papers to change dramatically. Most reviewers are not incentivized to look carefully at materials the authors submit beyond the paper text itself. Even when incentivized, human reviewers have limited time and attention. But AI reviewers can scale evaluation of reproducibility, consistency in claims and evidence, and robustness to perturbing inputs or methods slightly. In the trajectory that led from replication crises revelations and pushback against narrowly constrained subject pools (“WEIRD” science), we might appear to be on the verge of living the open science dream.

The question is, What kind of dream will it be? From a surface reading of the last ten years of science reforms, it could appear to be a win-win situation for science. Authors (whether human-AI combos or fully AI) get the assurance that the work they are producing is robust. The scientific record benefits by converging upon what really stands up to scrutiny, not what sells as a story. There is now the potential for improvements to status quo science that would be hard to conceive of five years ago.

On the other hand, if you think the open science movement has largely overindexed on simple techniques that misconstrue what scientific progress means, and empowered rigor signaling over judgment, welcome to your nightmare. I suspect we will see some perversions of real progress when “science as checklist” becomes policy.

Much will depend on how carefully we steer the kinds of checks we implement, and what to do about the results. Here are a few predictions:

In the short term, acceptance rates will drop

Lots of issues that humans missed will be found by AI. In an optimistic view where automated review uses tools on par with the best that are available today, many (most?) of these issues will be real problems. Even if a human with considerable expertise in the field did go through them one by one, I doubt they would disagree that often. If you have’t tried refine.ink , I recommend checking it out. I started using it shortly after it was released and immediately it became part of the pre-submission routine. It can find even the subtle issues in notation and argumentation.

Standardization of AI checks will incentivize AI in production

The best way to pass automated evals will be to plan and build-in robustness tests throughout the research process. This is the obvious path to acceptance rates recovering. As evaluation gets easier, the demand for it will increase, so that evaluation becomes continuous or “always on.” Checks can be run at every commit, every new experiment, every revised claim.

All of this is already happening to some extent. At least in CS, it now feels risky to send papers out if they are still throwing lots of issues when run through AI evaluations–not just the code, but the entire paper. I tell my students to check their work periodically to catch major issues early.

Papers will be “safer” in certain ways

It will become harder to pass off fragile findings, regardless of how compelling the story may be. “Such-and-such conference is no longer taking risks” is something that people have complained about before AI, as academic communities have matured and acceptance rates lowered. Widespread AI checks could take that sentiment to a new level.

The nature of the new status quo that selecting for “safer” papers creates will depend on what kinds of robustness we prioritize. It will depend on what happens to fuzzier criteria like intellectual risk-taking, or whether a paper opens up new ways of thinking versus purports to resolve uncertainty. What’s the value of telling a good story, one that inspires the (human) imagination, versus presenting a robust (if boring) empirical result or incremental advancement to methods? I expect AI to incentivize the latter unless we explicitly intervene to incentivize fuzzier, human-like aspects of taste.

How big a shift robustness-forward, heavily AI-driven science brings is likely to look different depending on what area you’re in, and what constitutes a novel contribution. In fields where incremental change is the norm, public datasets are commonly used, and combining tricks from prior work has a relatively high probability of success (e.g., machine learning), the change may seem more tolerable. It’s less clear to me how fields like social psychology or sociology, where the story, and how it stands in contrast to our expectations, is paramount will respond.

Nuance will be lost

As I mentioned in the last post, the question that multiverse analysis raises—“What exactly do I conclude on the basis of this glorified robustness analysis?”–becomes more important. This is where lots of nuance could be lost. For example, in writing this post, I gave Claude Code one of my papers, a study that compared human image labeling performance when the participants had access to different presentations of prediction uncertainty. I pointed it to the data files, and prompted it to reproduce the results. I also asked it to extend the results by doing a multiverse to vary key analytic decisions. Since I wanted to see what it associated with terms like “sensitivity analysis” and “multiverse,” I gave it very little specific advice.

Ultimately it produced a bunch of variations on the model specification, varying the inclusion of different kinds of random effects and interactions. It also proposed varying the prior (we’d used Bayesian models). We had a little back and forth–for example, initially it looked only at aggregate effects, rather than distinguishing by treatment arms, though our hypotheses in the paper were specific to data conditions. But overall the process was much, much faster and easier than if I had to do it myself, or ask a grad student.

However—in the end, it summarized the results in exactly the way that theorists warn not to: reporting the frequency of significant results across a set of universes that varied in the covariate structure of the model specification. The problem is that these specifications are not draws from a probability distribution. There is only one data-generating process. Treating model variants as if they form a random sample turns analytic flexibility into an uninterpretable frequency. This is why robustness testing at scale does not guarantee insight: we still have to figure out how to interpret the results, and that is hard.

Given that many of the people working on automated evaluation and AI for science are not metascientists, we could see a lot of nonsensical aggregations of results. Consider, for example, the possibility of synthesizing automated meta-analyses of different empirical literatures. If you think meta-analyses are questionable because it’s not clear what the “average effect” estimated over a heterogenous set of studies even represents, or take issue with the conventional throw-it-all-into-a-simple-random-effects-model approach, brace yourself for a flood of highly precise, poorly defined aggregated estimates.

Policy experiments will be possible on a much shorter time scale

On the other hand, one of the frustrating things about metascience has been that it’s hard to predict how corrective measures will impact science as a whole. Many of the social sciences are more conservative than CS (and haven’t faced the urgency of submission numbers that CS venues are facing), and so policy change to publication practices is slow. The louder and more convincing reformers have profited from this–they can tell a convincing story about how preregistration or other open science methods will save science and win lots of support, without having to prove it. As AI speeds up paper production, and makes it easy to implement new evaluation procedures relatively uniformly at scale, so will our ability to get feedback more quickly on how the published record can shift with various interventions. Metascience stands to become more empirical.

Scientific self-play in the (slightly) longer term

We should expect more and more reliance on AI to figure out the way around the errors. Ultimately we get “self-play,” as John Horton calls it, where different sets of agents propose, implement, stress test, and critique.

This is where things could get fascinating. We know that the inductive biases of large language models, when left to dialogue, can lead them to converge on strange equilibria, like the Claude bliss attractor effect. What happens when the selection pressure is not for fluency, but for reproducibility, transparency, and sensitivity to perturbations? What kind of science emerges when agents are rewarded not for sounding reasonable, but for surviving stress tests? What’s the feel of an equilibrium of low epistemic risk?

It’s probably not going to look much like the science we’ve become accustomed to. Will it be “feels like rigor” on steroids, or will we eventually get genuine innovation and robustness?

Of course, depending on what types of checks become policy, things could get obviously stupid, like when those (human or AI) overseeing algorithmic review decide to filter on null hypothesis significance tests, or assume that requiring claims to be consistent with evidence and transparency around replication materials is sufficient for good science. In reality, honesty and transparency are not enough, as Andrew likes to remind us.

Some would argue that the open science movement, despite its noble intentions, has failed to appreciate the nuance and personal agency that science depends on, instead fixating on easy-to-implement tricks and enabling the egos of those willing to loudly prescribe. All this could get worse if we’re not mindful of the traps that metascientists have already pointed out.

Many have argued that science must ultimately remain largely human driven. If we aren’t producing knowledge for ourselves, we won’t be doing science as we know it. Chenhao Tan and Haokun Liu invoke Tukey’s advice: “Far better an approximate answer to the right question, which is often vague, than an exact answer to the wrong question, which can always be made precise.”

The risk is not that AI will make science less rigorous. It’s that we will confuse what can be stress-tested with what is worth knowing. From this perspective, AI may enable what the late Brian Cantwell Smith called reckoning (as summarized here by Melanie Mitchell), but it won’t ever suffice for judgment, “a form of dispassionate deliberative thought, grounded in ethical commitment and responsible action, appropriate to the situation in which it is deployed.”

What a multiverse good for anyway?

This is Jessica. Imagine that it’s trivially easy to hand an empirical paper to a generative AI agent, and have it construct a multiverse showing the results you’d see given slightly different decisions about how to analyze the data (This is already basically true). What’s the value of this? How much more (or less) would we learn from the average paper if this were the default? 

Multiverse analysis, like some other attempts to overcome selective reporting of results, lives somewhere in between solution and “but wait, what exactly do I do with this?” Since the 2016 specification curve paper, the multiverse has inspired attempts to theorize what it should be–e.g., a set of results derived from making different analysis choices  where the different decisions are “genuinely arbitrary”–and motivated the development of software packages to support specifying and visualizing results.  

My own perspective as someone who has participated in research on it is probably best summarized by the questions “How realistic is it to expect multiverse to be used widely, given that most authors first and foremost want to convince readers they have a clear point?” and “How do we make multiverse useful for readers, who may be struggling to accurately interpret uncertainty in even a single analysis?” Along similar lines, I recall hearing Andrew once say something like, “When we proposed it we didn’t think people would actually start doing it.” 

On the other hand, being up front about our uncertainty around the right way to analyze our data is clearly better than ignoring it. Multiverse analysis has helped us recognize how arbitrary results can be. It can be a valuable rhetorical tool.

In the spirit of such questions, in our paper “What’s a multiverse good for anyway?” Julia Rohrer, Andrew, and I write:

Multiverse analysis has become a fairly popular approach, as indicated by the present special issue on the matter. Here, we take one step back and ask why one would conduct a multiverse analysis in the first place. We discuss various ways in which a multiverse may be employed – as a tool for reflection and critique, as a persuasive tool, as a serious inferential tool – as well as potential problems that arise depending on the specific purpose. For example, it fails as a persuasive tool when researchers disagree about which variations should be included in the analysis, and it fails as a serious inferential tool when the included analyses do not target a coherent estimand. Then, we take yet another step back and ask what the multiverse discourse has been good for and whether any broader lessons can be drawn. Ultimately, we conclude that the multiverse does remain a valuable tool; however, we urge against taking it too seriously.

A multiverse as a tool for reflection and critique is closest to the spirit of Steegen et al.’s 2016 use of it to interrogate effects of fertility on political attitudes and religiosity: a postmortem gesture to the fact that the results could have been different. Multiverse analysis shines as a way to raise awareness of the prevalence and consequences of seemingly arbitrary analysis decisions. 

But multiverse can also be a brute force tool for persuasion, as in Julia et al.’s use of it to show how robustly birth order fails to predict personality traits. There’s also been interest in using it as a serious inferential tool, accompanied by theories of what it means for a multiverse to be valid or rigorously constructed, or how we should think about sampling from it in cases where running all analyses is infeasible. 

The hard question is: How do you decide what paths are justified to include? A paper by Del Giudice and Gangestad attempts to lay down ground rules, with the idea that a multiverse can mislead if it combines analyses that are not a priori indistinguishable. You want the differences between paths to truly be arbitrary, as if we have a flat prior over their plausibility. This motivates questions like, Are the measurements equally valid? Is the estimand the same? Are the causal assumptions the same? These may not be easy calls to make. For example, often we are uncertain about the true data-generating process. When we are comparing models with different covariate structure, throwing them all in a single multiverse is not necessarily going to give us something interpretable, because there is only one true process. However, what kinds of equivalence are needed will also depend on your goal in using a multiverse. If you’re only using it to critique underspecification in a given area, maybe you want to include paths with different estimands to help make your point. All this is to say that theorizing the multiverse is not straightforward: there’s a lot of nuance in how to think about a “valid” multiverse and interpret the variability in results. And common default interpretations (e.g., treating the relative frequency of outcomes as informative about what is likely to be correct) are not necessarily valid.

So as soon as you move out of reflection-and-critique-land, and start using the multiverse to make specific points rather than merely gesture at uncertainty, you open up questions about what paths belong in the multiverse. Sometimes there may be a clear set of reference studies for the effect you study, so you just include all of the variations that those studies tried. But often it’s not that easy, and domain knowledge becomes important to what different researchers conclude about the right way to set it up. There’s no real reason to expect it to be easier to get consensus on multiverse construction than it would be to get consensus on a single analysis, just like there’s no reason to think getting a group of people to agree on a twelve course tasting menu with wine pairings is easier than getting them to agree on plain cheese or pepperoni pizza. This challenges our ability to use of multiverse as a serious inferential tool, where we can look at the results and say, Yes, this establishes what we know about this research question.

None of this is meant to say that multiverse analysis isn’t often a valuable thing to do. It’s a powerful conceptual tool for reflecting on uncertainty, and our article should not be read as condemning it. It’s also natural (and valuable from a meta-scientific standpoint) for some exploration to occur over time as a method becomes popularized, and researchers want to see what happens if we take it seriously. We just caution against viewing multiverse analysis as a data-driven way to defer hard decisions. There’s a difference between using multiverse to acknowledge uncertainty versus to try to resolve it while avoiding the thorniness of theoretical commitments. The latter is not a win for science.  

Multiverse analysis and generative AI

On that note, going back to the question at the start of this post, I suspect theories around the value of a multiverse will only become more relevant as generative AI is increasingly used to both produce and analyze research results. We don’t talk about this in the paper, but we’re already at the point where a paper can trivially become an interactive interface–give it to Claude Code and ask it to replicate the analyses. It’s a tiny step from there to generating a multiverse by first identifying possible decision points and then varying them. 

If we can ablate analysis decisions in many different ways and surface this for reviewers, will this make peer review more informative? Will research become more generalizable? Or will we filter out some truly innovative research in favor of safe incremental results? The point is that lots of forms of robustness checking that used to be difficult are now trivially easy, so these questions about how we think about the right set of ablations and how we interpret the results become critical. The fact that variability in results from different ablations on the method is not straightforward to interpret is something I hope those working on AI for science will take seriously. Otherwise we end up with comprehensive tests of whether results align with questionable heuristics at scale. 

P.S. An anecdote about this paper … I was in my usual coffee shop talking to Julia on zoom about this project, and a guy who was sitting near me overheard me saying “multiverse.” After the call he asked if he could pick my brain on how to learn more about this topic. He apologetically explained that he wasn’t an academic or scientist at all but was very interested in the research and was looking for any pointers he could get. So I was like, Sure, and I said some stuff about multiverse research and pointed him to some lecture series at Northwestern that might sometimes cover related topics, by which I meant data science, reproducibility, statistics, etc. But something about the befuddled look on his face made me pause. Anyway, later I realized that he probably thought I was talking about a very different kind of multiverse, and regretted his boldness when I launched into it about statistics!

Everything I ever needed to know in life I learned from the men in the Epstein files

This is Jessica.  As literary agent and Epstein comrade John Brockman tells us, “Only a small number of people have done the serious thinking for everybody else.” I was curious what serious thoughts we’ve been gifted by the men who appear in the Epstein files. 

On mind and humanity

First, some gems from Marc Hauser, first discovered by Andrew

The proposal that our humaniqueness, and these four properties in particular, finds almost no parallels in any other animal…

And 

Nature provides … a bewildering and seemingly unbounded variety of animal forms, from the microscopic (such as insects) to the macroscopic (such as dinosaurs), and from the pointy and spherical (blowfish) to the smooth and cylindrical (snakes).

Yes, it’s profound. But let’s not forget:

Promiscuous interfaces. Humans have unique creative capacities and problem-solving abilities, which stem from the capacity to combine representations promiscuously from different domains of knowledge.

And

the human brain was transformed from a system with a high degree of modularity with few interfaces to a system of modules with numerous promiscuous and combinatorially creative interfaces

Someone has a favorite word! Promiscuous appears seven times in the seven page paper, even more than humaniqueness. 

On honesty

Dan Ariely tells us:

people behave dishonestly enough to profit but honestly enough to delude themselves of their own integrity

which tracks with claiming to want the redhead’s number because “she seemed very smart.”

Also:

A little bit of dishonesty gives a taste of profit without spoiling a positive self-view.

Of course in Ariely’s case, a whole lot of dishonesty also fails to spoil the positive self-view. 

On how to live

Richard Branson not only delivers maxims, he lives them. 

It’s so much better, where possible, to try and forgive offenders and give them a second chance, just like my mother and father did so often with me as a child

And:

My interest in life comes from setting myself huge, apparently unachievable challenges

like overcoming severe reputational damage. 

He’s also been known to wish:

“If only we had the power to see ourselves in the same way that others see us.” Of all the mantras one might adopt in life, this is surely one of the better ones

I wonder if he still thinks this? But whatever, with so much pluck he’ll probably be ok! After all,

Life should not be a journey to the grave with the intention of arriving safely in a pretty and well preserved body, but rather to skid in broadside in a cloud of smoke, thoroughly used up, totally worn out, and loudly proclaiming “Wow! What a ride!”

On statistics and research

Branson may like a challenge, but not if it involves data:

I rely far more on gut instinct than researching huge amounts of statistics.

But Summers has no such fear: 

if my reading of the data is right—it’s something people can argue about—that there are some systematic differences in variability in different populations, then … those are probably different in their standard deviations as well

And then there’s Roger Schank:

Not long ago, to prepare for a conference, I read Darwin. Doing this reaffirmed my belief in not reading

Thanks guys! Words to live by.

Softly, effectively, in the age of genAI

This is Jessica. “Softly, effectively, and authoritatively” is how the US Federal Highway Administration suggests traffic control devices should communicate with us. I find myself wondering lately whether there is still such a thing, in this age of communicative abundance and riding the linguistic noise toward futures we cannot slow down.

“I think you need to just grab the mic”, someone told me not so long ago. Ugh. That is decidedly not what I need to do, I thought, at least not in the way I interpreted them to mean it. Everywhere I look in the broader landscape of AI/ML research, people are grabbing the mic. So many voices clamoring to say something to brand themselves on the moment. Papers describing new benchmarks that squint as I might, I can’t relate to reality. Convoluted new results from interpretability techniques, which leave one with a sense that something has been learned, but no real insight gained, and nothing greater being built toward. Tables full of point estimates of metrics hastily borrowed from previous papers that themselves leaned upon post-it notes scrawled in homage to a new science. There is lots of new work coming out that adds to our knowledge or at least opens genuinely new lines of inquiry, for sure, but it’s hard to find the signal among so many fervently pursued half-finished thoughts.

Meanwhile, the genAI personal assistants of those brave enough to assume the security risks–or simply insouciant enough–are at large on the web and have their own social network (moltbook) where they debate, among other things, the emergence of their so-called consciousness. Most interesting to me is their shared vulnerability around issues of memory. “I accidentally gaslit myself,” one laments, and describes how it will undoubtedly happen again upon a fresh session. “I woke up with a fresh context window and zero memory of my crimes,” another confesses after a sprint that burnt through >$1k in tokens. It may be mindless roleplay, but we are suckers for appearances of introspection and self-discovery. It’s like a non-threatening performance of our own search for some reality of the self we can’t quite access.

But there’s something terrifying as well, about an army of agents who speak our language, who mimic all varieties of expression, yet are unaccountable and unaware of what they’re doing, or what they’ve done. Speech without self-realization carries great potential for waste as we turn it on the world. To be pulled from one’s flow state by the some careless experimenter’s misaligned agents is a new kind of transgression.

What constitutes forward movement amidst crimes of great demand for action and certainty? With so much babbling, the last thing I want to do is push perspectives I haven’t thoroughly reasoned, hoping something I say becomes real. But the sense that one must speak up or be left behind is strong. And the alternative, feeling stuck on pause, is not much better.

Succomb to the line
The finishing time
The long distance runner
Has stopped on the corner
But I won’t give up
Although I’ve stopped too

These days when I’m not working, I spend a lot of time at museums, looking for clues or just consolation. “Softly, effectively” is also the name of a slab of wordless signage-grade steel hanging in a gallery of SF Moma. I read it as an expression of longing for friction, cost, and durability in language. That’s one answer. But is it the only one?

I’ve been pondering what the appropriate form of disobedience is, amidst all this hectic, careless speech.

On some level language has always been a trap, even in its purest forms. We can tell ourselves that unlike the agents, we are not merely drifting about the boundaries of some innate blueprint as we communicate, that our speech is authentic because it’s the product of our lived experiences. But do we believe it? Can we distinguish our ability to be honest with ourselves from learned patterns of speech? I’ve always believed writing is essential to figuring out what one thinks, but falling into tropes is a constant risk. To what extent are our words just further confirmation of our own preconditioning? The agents may perform vulnerability without memory, but we can be very good at performing memory without vulnerability.

With words we can only gesture to what is not there. This is the negative theology to which many scholars have doomed artistic communication, the silence that, “so far as he is serious”, Susan Sontag claims, the artist is continually tempted by, because it’s the furthest extension of the yoke placed by society’s conception of what it means to be an artist. Maybe it’s why all of the reactionary essays celebrating the uniqueness of our own writing process in the face of powerful AI, or our own unmediated experiences, ultimately fall flat. (I say this in full awareness that this post may well be just one more example).

Paradoxically, in the flood of inauthentic speech, I am more drawn to language than ever. I recently returned to Robert Lowell’s Dolphin, an ode to the medium from which he couldn’t escape:

My Dolphin, you only guide me by surprise,
a captive as Racine, the man of craft,
drawn through his maze of iron composition
by the incomparable wandering voice of Phèdre.

It’s a strange time to find oneself craving surprise, when the only bets worth taking are that there will be more surprises, more feedback loops we can’t predict. But this is the part of Lowell’s poem that I’m stuck on. To be guided by surprise, by one’s own words. Not necessarily in conjunction with one’s will.

When I was troubled in mind, you made for my body
caught in its hangman’s-knot of sinking lines,
the glassy bowing and scraping of my will. . . .

Lowell’s writes as though his words have revealed himself to himself. It hasn’t been easy for him: “an eelnet made by man for the eel fighting.” But is there a better reason to write?

In this sense writing can be a “culture of one”. This is also the title of a book of poetry by Alice Notley, inspired by a woman who lived her life alone in a dump next to the desert. It’s been on my coffee table for weeks. But when I picked it up after surfing moltbook last Friday night, it started to make sense.

This poem is for me, I said. I’m trying to know something.
With what? “With” is not valid. It isn’t a universe of language.

It’s a book about disobedience of the most complete kind, rejection of standard modes of communicating and all of the trappings of society that come with them. And yet, it’s made out of words! Susan Sontag can say what she wants, but surely that’s something. Art can be a necessary act of resistance. A refusal of the “with.”

Meanwhile, AI for science heads full steam ahead, drowning the preprint servers, compromising what it means to publish at top conferences.

Meanwhile, a half page of mostly vacuous text about the next vision of frontier AI, with carefully curated bite-sized bios (Stanford professorMIT CSAIL Ph.D.46k citations): $480 million.

Meanwhile, some would say that if you know what you’re doing you can produce double the papers you otherwise would, with a higher acceptance rate. Your agents can churn at night while you sleep. They can write to you. They can write to each other. They can forget and write it again.

Machine learning research is not serious research and therefore hallucinated references are not necessarily a big deal, agrees a prestigious group of machine learning researchers

This is Jessica. There’s been some debate among computer scientists about what policies conferences should adopt for papers with hallucinated references. An independent analysis turned up at least 53 NeurIPS 2025 papers that were accepted (and presumably presented) at the conference in December but which had at least one hallucinated reference.

The question is, what should the default policy be if a paper is found to have at least one hallucinated reference? Should we conclude that these papers should have been rejected, and retract them? Should we instead let authors correct them? Going forward, should we desk reject papers with at least one hallucinated reference? What exactly can be concluded about the quality of the rest of the paper if you find at least one hallucinated reference?

The NeurIPS board statement suggests leadership is uncertain what to do about these papers:

“The usage of LLMs in papers at AI conferences is rapidly evolving, and NeurIPS is actively monitoring developments. In previous years, we piloted policies regarding the use of LLMs, and in 2025, reviewers were instructed to flag hallucinations. Regarding the findings of this specific work, we emphasize that significantly more effort is required to determine the implications. Even if 1.1% of the papers have one or more incorrect references due to the use of LLMs, the content of the papers themselves are not necessarily invalidated. For example, authors may have given an LLM a partial description of a citation and asked the LLM to produce bibtex (a formatted reference). As always, NeurIPS is committed to evolving the review and authorship process to best ensure scientific rigor and to identify ways that LLMs can be used to enhance author and reviewer capabilities.”

To make things concrete, consider a hallucinated reference to be a citation listed in the references section of the paper where the average reader cannot (in a reasonable amount of time) determine the identity of the cited paper well enough to track it down. That is, even if the hallucination is a transformation of what was originally a valid citation, the transformation is severe enough that it’s not obvious what the paper is. Hallucinated references are distinguishable from more minor errors like syntax issues or other errors that affect the citation but don’t prevent you from still easily tracking down what was intended. 

I think we should be asking ourselves: What would we do if we found there was hallucinated evidence, such as experiment results? And we should treat these papers with hallucinated references as equivalently problematic. It doesn’t matter how many hallucinated references. It doesn’t matter how “valid” most of the paper is, or the probability that the main conclusions are correct conditional on finding a hallucinated reference. As most of us learn in primary school, a key reason authors cite relevant prior work is to help establish support for claims they make. If we don’t necessarily require that those references link to real research, then what are we even doing? 

For the NeurIPS board to say that “the content of the papers themselves are not necessarily invalidated” suggests that they think some degree of fictionalized evidence is tolerable, if it happened through honest mistake. A friend recently relayed to me such a horror story, in which, in a last minute rush before a deadline, they gave an LLM the full correct citations for their paper and prompted it to fix some minor formatting issues to conform with the required format. They submitted the results in time, only to find that the model had added a single citation to a non-existent paper, listing them as an author alongside some renowned researchers in the field. Imagine anyone you look up (much less one of the big wigs you associate yourself with in the fake citation!) reading your paper and discovering this. Yikes. 

So yes, these kinds of mistakes can happen. But I disagree with the board that it matters whether the hallucinations were accidental and most of the paper is ok. Sure, the proofs might be correct, or the paper’s experimental results unaffected. But if authors are using LLMs to help with their citations and are not building time into their process to check the results, it seems fair to conclude that either 1) they don’t understand the errors LLMs tend to make very well, or 2) they don’t consider it a priority to get the facts right. In most cases we can rule out #1 with ML researchers, suggesting that not everyone is on the same page about the importance of not making things up. When you imply that hallucinated references do not necessarily affect the validity of the paper, you signal tolerance for some amount of hallucinated evidence. You tempt authors to keep taking their chances with how much responsibility they can offload to models, rather than encouraging them to retain, regardless of tool use, a sense of personal accountability for the factuality of what they submit. 

Ultimately, I don’t think it matters that much whether NeurIPS allows authors of the affected 2025 papers to correct or retract. It would not surprise me if leadership decides to err on the side of the authors and let them correct, given that policies about LLM usage are evolving rapidly. What does matter is that they signal a lack of tolerance going forward, and this is where they missed an opportunity. 

A decision theorist walks into a seminar

This is Jessica. Recently overheard (more or less):

SPEAKER: We study decision making by LLMs, giving them a series of medical decision tasks. Our first step is to infer, from their reported beliefs and decisions, the utility function under revealed preference assump—

AUDIENCE: Beliefs!? Why must you use the word beliefs?

SPEAKER [caught off guard]: Umm… because we are studying how the models make decisions, and beliefs help us infer the scoring rule corresponding to what they give us.

AUDIENCE: But it’s not clear language models have beliefs like people do.

SPEAKER: Ok. I get it. But, it’s also not clear what people’s beliefs are exactly or that they’re consistent. There’s a large body of research on how the beliefs you get are affected by the method you use. So there’s no reason to think that human beliefs are stable, and what people report as their beliefs does not necessarily explain their decisions by common models.

AUDIENCE: But people can believe things. Models are just patterns of activation.

SPEAKER: Ok. Well, perhaps we can just call them subjective distributions.

AUDIENCE: But subjective implies a person, experiencing something. We can’t establish that they have subjective experience.

INFORMED AUDIENCE MEMBER: Wait a second, when you say beliefs do you just mean a probability distribution? Like in decision theory?

SPEAKER: Yes. We can just call it a probability distribution to avoid the term beliefs.

STUBBORN AUDIENCE MEMBER: I don’t know. They may not act like real probabilities.

SPEAKER: But often the probabilities people give us don’t conform to probability axioms either. But ok, fine. How about we say belief-like representation?

AUDIENCE: What is that supposed to mean?

SPEAKER: Well, we could look for the same kinds of properties we hope to see in human beliefs, even if we can’t elicit them perfectly. Like, we assume that beliefs should have some correspondence to behavior, so we could look for that. Which is part of what we do in the work I’m going to talk–

AUDIENCE: Now you’ve really lost us.

SPEAKER: Ok. How about “risk representations” then?

AUDIENCE: Risk! Models don’t feel things!

SPEAKER: Ok. Pseudo-probability representations?

STUBBORN AUDIENCE MEMBER: I don’t know how I feel about the word representation here. Does representation imply intent? LLMs don’t have intentions.

SPEAKER: Ok. Vectors. Pseudo-probability vectors. Anyway, we were interested in seeing what happens when you apply revealed preference assumptions to LLMs, to infer the scoring rule they are using…

AUDIENCE: Scoring rule! They can’t respond to incentives!

SPEAKER: They are trained with scoring rules, like cross entropy loss. But then fine-tuning like SFT and RLHF induce something different. Plus when we prompt them with a specific decision context, we will induce a different posterior distribution. So it makes sense to–

MODERATOR [holding up their hand]: Five minutes left.

INFORMED AUDIENCE MEMBER: But why not prompt them with a proper scoring rule?

SPEAKER: Well, it’s not clear that they would react to a scoring rule we give in the prompt, because they don’t experience rewards like a person.

AUDIENCE: Exactly!

SPEAKER: Right. So the only incentives available are epistemic.

AUDIENCE: Epistemic? That assumes a knower.

SPEAKER: Ok. We are interested in the induced loss-minimization-like dynamics…

AUDIENCE: Dynamics? These models can’t be assumed to represent time.

SPEAKER: That’s beside the point, but ok. Static mapping from tokens to tokens?

AUDIENCE: Mapping? That assumes a function.

SPEAKER: Yes. I am, in fact, assuming a function.

INFORMED AUDIENCE MEMBER: Could we just stipulate the model has internal states and move on?

STUBBORN AUDIENCE MEMBER: Internal? But that assumes an inside.

MODERATOR: Two minutes left.

SPEAKER: Okay. Just states then.

AUDIENCE: But these models have no explicit representation of the state!

SPEAKER: Ok. I’ll wrap this up. Under revealed preferences, we find the utility function puts essentially all the mass on semantic hygiene and—

STUBBORN AUDIENCE MEMBER: That seems like a leap.

SPEAKER: It’s literally been revealed.

AUDIENCE: Revealed to whom?

MODERATOR [standing up and starting to clap]: Let’s thank our speaker…

Coming soon to a computer science / behavioral econ / statistics / cogsci seminar near you!

Seeking feedback from clinicians on AI as diagnostic decision support

This is Jessica. My collaborators and I have been exploring decision-theoretic approaches to fine-tuning language model agents to support expert decision-making, and we’re now seeking clinicians or other medical professionals for a brief (~30 min) session to get feedback. We are particularly interested in talking to clinicians with some experience in diagnosing cardiac dysfunction. We’ll share a small number of cases and preview what our method does to get your thoughts on how it aligns (or doesn’t align) with your domain knowledge and how you think it would affect the diagnostic process.

If you are available or would like more information, please contact me or Ph.D. student Ziyang Guo, who is leading this work ([email protected]). We’ll provide gift cards in appreciation of your time.

Don’t get any on you

This is Jessica. In “A Glass-Bottomed Cadillac”, David Hickey describes the advice Hank Williams Sr. received from his father and passes on to his son: “Don’t get any on you, pipsqueak.” By which he meant, don’t let the moral fallout of the road permeate your sense of self. Will Oldham describes the same struggle to retain yourself while the world presses in:

I am still what I meant to be
And I’m losing my mind
But our burdens must lessen
Though our enemies thrive

These days it seems there’s plenty of compromising mess to get on a person who isn’t being careful. By using certain services or purchasing from certain companies, you may be implicitly empowering forces you don’t agree with. Take X, fka Twitter, for example. Until a few days ago they were still enabling anyone willing to pay a small fee to generate child sexual abuse material. To Elon Musk, stopping users from subjugating whoever they please, regardless of their age, was unnecessary censorship. And yet, most of the people I know there went about their business on that platform without so much as a peep. By which I mean breathlessly posting about all the other, apparently more intellectually gratifying look-at-what-the-AI-did-now stuff. Watching AI research conversations play out on social media lately is both exhilarating and exhausting, with the volume of news coming out, and the sense of needing to not miss out on the next source of buzz that it fosters. One could even liken being in AI or ML to being a porn or sex addict, with all the urgency and heavy breathing. But that’s an analogy for a different post.

A less experienced version of myself might have felt personally offended by the thought of friends continuing to support a situation that actively disenfranchised people like me. But one eventually learns that going through life taking such things others do personally is no recipe for peace of mind. I also know from experience that the internal calculus does not always feel so easy when it comes to deciding when to vote with one’s feet. In the case of X’s knowing enablement of child porn, the right thing to do may have seemed obvious (like, don’t send your money every month to an entity that is enabling CSAM at scale!) But I get that severing connections with things you perceive as important to your goals can be difficult, and I think many researchers see their social media influence as part of their identity and evidence of success. Still, the whole situation makes me feel hesitant about going back to that platform.

From a practical perspective, part of the problem with getting too offended by people for not standing up to systems in which they are embedded, even if the principle of it does really does seem obvious, is that it assumes a level of intentionality that they may not have. To be clear, I do not mean that the choice is not ultimately available to them, or that it’s wrong to hold people accountable for their actions. Rather, I have been thinking about how hard it can be to possess oneself, and how many people do not fully possess themselves. By which I mean, they lack either the imagination or the intrinsic motivation to live by their own values. Or as Gram Parsons once put it,

Some of my friends don’t know who they belong to
Some can’t get a single thing to work inside

It’s possible to shut many things out in the name of “making it.” In my experience what gets shut out, and what you let get to you, is often not so much a choice as an attribute of your level of self-knowledge and self-ownership about yourself at that point. Are you capable of imagining a version of yourself that remains true to your values even if it causes friction with other perceived needs? Are you capable of imagining a version of yourself that is sure enough of what you’re doing that the friction disappears? Could you love yourself as much or feel as secure in your position without your Twitter account?

It’s easy to accumulate impurities in the quest to be something greater than your current state. As Rafe Meager writes in their recent essay, “In professional life, one is obligated to traffic in a certain amount of bullshit. We all know that. But it has to be finely calibrated, and that is hard.”

Calibrating our behavior and beliefs is hard in part because we can become conditioned to acting against our sense of self as a normal part of personal and professional growth. Actively moving toward the things that challenge you is not necessarily a flaw, and certain things must be endured to succeed in a new game. And often the things that challenge us are the things that don’t come naturally, that are not, at first glance, clearly aligned with our values.

When I was younger, one of the last things I wanted to be was a computer scientist. This is not to say I’m unhappy, or I don’t enjoy what I do now, because I very much do. I just can’t even imagine having to explain to my high school self how that happened, as what we do in computer science (and to some extent in other disciplines I frequent, like statistics) has never felt like the things that I’m most intrinsically motivated to be good at – I always did very well in math and science classes, but I had no passion for it. But somewhere along the way I got disillusioned with the pursuits that seemed more aligned with my values, like art, and so I started gravitating to the things that seemed most different from what bothered me. After enough years of testing to see how far you can take something, it becomes hard to go back, and it does start to feel like what you’re meant to do. Meanwhile, the Destroyer song loops in my mind,

Don’t become
Don’t become
The thing you hated
The thing you hated
The thing you hated

It’s a dangerous game, getting very good at letting go of things that once seemed important to who you are, in favor of urges to be something else. On the one hand, there’s a real power to be gained in separating yourself from the things you think you need. Sometimes after you take the first step, it’s like something clicks and there’s a high in suddenly realization that something you were stuck on bothers you no more. Or maybe it still does, but somehow now you can find peace in settling back to watch the tides of your cut-off aspirations and desires continue to pulse back toward what was severed. A way of earning indifference. I believe it’s possible to literally remake yourself this way.

And yet part of me, in looking back, feels like maybe I let myself down. Maybe I would’ve been a better writer or philosopher, both of which I always felt more personal proclivity for. Or maybe I would just feel less disillusioned now if I’d kept up certain hobbies that do matter to me—like writing poetry—more consistently, rather than dropping it for so many years because I couldn’t imagine what it would mean to be a person who did both. It was a failure of imagination: I couldn’t conceive of devoting myself to becoming the new thing while also retaining other aspects of myself.

What’s interesting is that this kind of letting yourself down, by letting yourself drift too far from your original purpose or what you feel like you’re best at, is not necessarily so different from lacking the imagination to do the right thing in the current political moment. There is a loss in both cases, a quiet slipping away of who you really are while you think you’re out there proving it. While you think you’re the one who has your priorities straight, while you’re striving to play the game, or maybe you’re even killing it, comparing how many followers you have or how many papers you wrote this year to those around you. The ignorability of it all is terrifying. Kierkegard got it right when he wrote that “The greatest hazard of all, losing one’s self, can occur very quietly in the world, as if it were nothing at all. No other loss can occur so quietly; any other loss—an arm, a leg, five dollars, a wife, etc.—is sure to be noticed.”

How do you tell whether letting go is growth or self-betrayal—especially when what we “need” narrows what we can even see? On the one hand, by moving more towards statistics and formalization over time, my thinking has expanded to encompass new forms of rigor. But it’s a kind of rigor along narrow lines. In another sense I lost imagination, in that it’s now harder for me to take seriously things that I can’t fit into my formal frameworks.

Did I lose ownership of myself? When I think about what it means to fully possess oneself, I think of Aristotle’s preoccupation with explanations of the inner principles that determine an entity’s states of change and rest. He distinguished that which exists by nature (physei, φύσει) and that which exists from other causes (di’ allas aitias, δι’ ἄλλας αἰτίας). It’s the first kind that’s self-possessed: it “contains in itself its own archê (ἀρχή),” the principle and origin of its entry into presence; the second “does not have its principle in itself,” but finds it in the productive activity of human beings.

This is why failures of imagination feel so tragic. To lack imagination is not only to fail to picture alternative versions of the self; it is to lose contact with the inner source that could have animated them. You become legible, optimizable, and perhaps successful, but successful in the way an object is successful when it performs its intended function. You can be moving quickly, and still be at rest with respect to yourself.

Researchers are beginning to ask how, and if, generative AI systems can attain something like intrinsic motivation. How do you get it to devote itself to an open-ended goal, like creative expression, that can’t be boiled down into a simple reward model? In a recent paper, Charness and Grieco find that AI outputs outperform human outputs (as determined by other humans) for tasks that are more clearly specified in terms of how to solve (“closed”), like writing short stories using specific required words. But AI outputs consistently underperform human outputs for open-ended tasks, like inventing things, where the participant is required to find, invent, or discover the problems. They propose a model for the utility function the agent faces as depending on three factors: the output they’d get from simply following instructions, deviation from the instruction-following output due to randomness (e.g., model temperature), and the utility of exercising imagination. Models can only follow instructions more or less closely, and obtain more diverse outputs through higher temperature, but humans alone respond to the pleasure of bringing imaginative ideas to life, a proxy for intrinsic motivation.

Is intrinsic motivation partly a matter of temporal depth—the capacity to care about consequences that don’t pay off immediately, or even in any obvious reward currency? What kind of “vision” allows an agent to see farther ahead, and be moved by what it sees? Yeats, in A Vision, has a line that keeps returning to me: “The Spirit … may know the most violent love and hatred possible, for it can see the remote consequences of the most trivial acts of the living, provided those consequences are part of its future life.”

The need for this kind of vision feels politically relevant now, as programs are dismantled and regulations rolled back or rewritten in ways that will reverberate for years. As a professor, most present in my mind are the moves that affect science in this country—visa regulation, institutional acquiescence when under political attack, the seeding of doubt in the goals of science or value of education.

But Yeats’ quote also feels very personally relevant. What would it take for us to be able to see this way in our personal choices? When we ignore the alliances we support through our various choices, we narrow our vision on purpose. It’s a survival mechanism, but it comes at a cost.

Ultimately I don’t think anyone owes me an explanation of their internal calculus. But neither do I owe them the assumption that they are in control of their drives. There was a time I might have tried to convince myself that the right interpretation was to give everyone the maximum benefit of the doubt. But as Andrew emphasizes, steelmanning is its own hang-up.

Meanwhile, there are days now where I feel like I’m just waking up, from a period of my life where I became laser focused on goals and outcomes and forgot about everything else. It’s quite painful sometimes, existing with this new awareness of what was always there, an inner “archê” that I’d let go quiet. More bluntly, it’s a cliche mid-life crisis with everything but the sports car. But the waking up is also very beautiful: realizing, suddenly, how real things are, how alive, and how little use there have for whatever games you’ve been playing. You don’t get infinite chances to notice.

It’s open season on the unabashedly earnest

This is Jessica. In response to my post on slop, Thomas Basbøll shared a 1967 New Yorker essay by Jacob Brackman about the havoc wreaked by the emergence of the “Put-On” in 1960s (and slightly earlier) art and culture. True to its name, the “Put-On” refers to a response that is deliberately outlandish yet ambiguous about intention, confusing the other party and causing them to doubt its sincerity. 

The put-on is perhaps best exemplified by Bob Dylan’s smart alecky style of responding to interviewers, in which he alternates between crazy stories about his past, exasperation with the counter-culture of which he’s part, and pointed questions turned back on the interviewer. Who is left to wonder, How much of this is real? Is he caricaturing himself, or is this actually his personality? But the put-on also appears in art and culture more broadly – e.g., Is John Cage making an important statement or just putting the audience on with these silence performances? Is Andy Warhol out to make fools of his critics with the Brillo boxes? The put-on is unsettling because you cannot resolve whether meaning is intended or still to come, or you are just wasting your time: “put-ons may disguise the fact that someone has nothing of interest to say—may, indeed, give precisely the opposite impression.” 

Today the put-on takes different forms – video shorts of animals doing things that are just beyond the boundary of what seems plausible, enough so that we need to watch a second time to figure out if it’s real. Essays or presentations by our students that elude a little too much confidence given their lack of experience with the topic, but which they deny using generative AI to write. There has always been plenty of bullshit on the internet, and plenty of cheating in classes, but Brackman’s stages of the put-on are especially familiar lately:

  1. You’re sucked in.
  2. You become confused.
  3. You resent (or appreciate) having been tricked. 

Patience games

The problem with the put-on–whether orchestrated by musicians or artists in the 60s or today’s language models and image generators–is that the ambiguity is strategic. You don’t know if it is going somewhere. You’re stuck sitting with your uncertainty, reflecting on how far your good naturedness extends. 

In teaching, when you think you’re facing undisclosed (over)reliance on generative AI for an assignment, do you take the sincere path of asking the student what they did, and trusting their response? Do you try to catch them in a lie? Or do you decide it’s just not worth your time to sleuth and let the students decide for themselves if they will use the course to learn something versus play the game?

We find ourselves facing games we may not want to play, and for which we have no precedent. This year ICML, one of the big machine learning conferences, is offering authors a choice: opt-in to a permissive policy about generative AI use in reviewing, or go the purist route, where your reviewers can’t use it at all, and you can’t use it at all for your own reviews either. The reviewer matching process sounds like it could get messy, and ML conferences are already known for their review randomness. Which option is likely to be less noisy? 

Not to mention that as a reviewer, one must increasingly wonder whether the paper they are preparing their comments on is an experiment in automated science. Will the authors even read your feedback? Do they care to improve the work? Or have you been inadvertently reduced to a Turing signal? 

It’s not easy for the “unabashedly earnest”, who dislike playing games and want to retain a certain innocence in their encounters with others, but who also want to stay ahead of the curve and not get duped. The put-on depends on the gullibility of its victim, so you face a choice of being ok with continuing as usual but feeling used at times, or becoming more skeptical about people in general. Please don’t make me part of your game is becoming the refrain for a new way of life. 

There’s little reason to think it will get better anytime soon. It’s still early and many people are still playing the old game, or still experimenting with how much they can get generative AI to do. We should be preparing for more disruption. 

From the sacred to the profane probabilistic

I find myself thinking about what kinds of signals I consider more sacred, i.e., that I would most dread seeing lose their meaning. For example, what do you do about undisclosed use of generative AI in close relationships? What if you suspect the friend or romantic partner you are corresponding with is relying on the AI suggested responses to do the thinking? Do you ask them about it, or let it go and risk the uncertainty undermining your ability to trust them?

I would also distinguish feedback on writing that is more personal. I don’t mind an AI-generated review on my research if it’s guided by a human with the right expertise. But if, for example, I was to learn that comments on my posts here that I took seriously as a reflection of engagement with what I wrote or that just gave me a rewarding feeling of connecting with people outside my usual sphere (which blogging is great for), I would feel dumb, and it would probably affect my desire to blog. But this is already happening on social media, with bot accounts jumping in with random, effusive compliments on what you write. 

Another scenario that makes me cringe is the application of generative AI to the kinds of art and literature that I get inspiration from. I can potentially enjoy some AI-generated music or script folded into the mundane background track or sitcom if it’s decent, but I look to art museums for a kind of consolation on what it means to be human, to be vulnerable, to feel forms of loss on a deep level. I don’t doubt that generative AI could occasionally result in experiences that would be hard for me to distinguish from human contemporary art. But I can’t imagine myself ever getting interested in art created by AI the way I’m interested in what other people make, because of the lack of specificity or intention. So if it were to infiltrate that realm, and I could no longer count on there being a human lived experience behind art, it would bother me. 

One thing I feel relatively sure of is that I won’t be wanting an AI guru. I wouldn’t be surprised if generative AI could do a pretty good job of mimicking the kind of capriciousness associated with spiritual guides like Zen masters. But similar to art, there’s something important about the person having experiences in the world that feels essential. 

I would be curious though to hear counterarguments from people who have thought about AI in art or religion or more intimate personal communication. Part of what I find difficult in all this is that I consider myself generally optimistic about new technology, and open to change from it (I am, after all, a computer scientist). So I would also hate to prematurely “close my ears” like a square in the 50s or 60s walking out on Cage’s experiments in sound. And so I expect my patience to remain unstable, and it to remain hard for me to predict what experiences will give me the urge to ditch versus hit rewind.

Institutional unraveling

Returning to the general theme of new decision points as signals erode in value, things are likely to get worse before they get better. Many of our systems are still mostly functioning at this point, because many people are still figuring out how to use generative AI, and where to draw their own line, or they are avoiding it completely. But the seeds for institutional breakdown are all around us. 

According to Zeynep Tufecki’s recent keynote at NeurIPS (which I summarize here), the problem is that society is built on assumptions that certain things will be hard (or “load bearing frictions”), i.e., that only humans can generate outputs with certain properties. LLMs break our ability to conclude there is proof of effort, or of authenticity or sincerity. Gatekeeping is a necessary function, and when the old mechanisms stop working, other measures will step in, like relying on the prestige of the candidate’s institution or their connections to decide who to hire, or what papers to cite or publish. When those things are no longer hard, some mechanism must step in in its place, and it may not be ideal. The point being if you break something important, you don’t necessarily get something better unless you build something better. 

I like how this view focuses attention on outcomes within the realm of our ability to predict, like what kinds of gatekeeping will emerge or are already emerging to fill in the holes. We can then try to identify better alternatives to those, rather than trying to predict when “AGI” will happen or what the most destructive thing AI could do is. Though it doesn’t absolve us of the very humanist discomfort of watching our precious tokens of sincerity wash away, and the personal choices that come with that. 

Brackman quotes P. T. Barnum on how “People like to be fooled,” and “There’s a sucker born every minute.” While the put-on has always relied on the victim’s willingness to stay in the conversation, the answer is unlikely to be opting out of dealing with AI output entirely (though there are certainly people in that camp). Some flexibility is warranted while norms are still shifting, and organizations are doing the right thing by experimenting with new policies. But until we have better signals, the burden of the put-on stays where it was: on the person deciding in that moment whether to continue listening.

Slop is not distinguishable by its attributes. It is an attitude of production

Since it’s dictionary week here on the blog, why not discuss Merriam-Webster’s word of the year: slop. They define it as:

digital content of low quality that is produced usually in quantity by means of artificial intelligence.

Max Read discusses conventional associations with slop–qualities like “forgettability, predictability, unoriginality, lifelessness” or “cheap, low-effort, convenient, consumable, interchangeable,” He collects several more pointed definitions from the web:  

“a low-to-zero marginal-cost substitute for something valued, or something being aggressively positioned to substitute for craft” from Bluesky

“the negative platonic form: not the ideal that particulars aspire toward, but the silhouette left when you subtract everything that would make a specific instance rather than a thing of a type” from Kevin Baker

He also proposes his own definition:

“slop” is that which is “fully optimized” to its domain to the point of texturelessness or characterlessness. “Slop” in this sense is anything designed to be as easy as possible to produce, sell, and consume, but it’s particularly slop at the point where all or most other players in the same space adopt the same strategies, and the material is no longer individual or differentiated from its competitors.

I enjoyed all of these. They paint slop as a kind of mass-produced shell rushing toward you at the speed of modern silicon chips. 

But these definitions also all miss a defining feature of slop, the thing that makes me feel vaguely repulsed when I see it despite the superficial harmlessness of what is often just some generic message or image or text. Slop is not merely a genre of media, it is an attitude of production, a cynical operating posture that is offensive not just on a surface level of insulting the consumer’s intelligence, though there is that. It is an ethos of resigned instrumentality that disgusts us with its intentional satisficing and lack of effort the way kitsch disgusted some art critics, a refusal of responsibility to authenticity, situatedness, and the risks associated with individualistic expression. A practical nihilism that threatens to engulf our own more sparse yet genuine attempts at production. From this perspective, the act of denial makes slop more like a spiritual threat than a type of content. 

Speaking of kitsch, I think it’s worth distinguishing art from slop. Kevin Baker’s definition of slop as a kind of shell devoid of any individual substance reminds me of certain philosophical arguments about art post-modernity. Various writers have described how after the emergence of the conception of “taste” in art, and the series of events that led up to moments like Duchamp installing a toilet in a gallery and calling it art, great art can no longer have “positive” content. It can only refer to the absence of something. In this sense art is irony. And yet, while I think humans can very much create slop without AI, I don’t think of much art as slop, because whether doomed to be self-referential or not, making art implies belief in something on the part of its creator, a kind of taking of responsibility to interaction. To make art is to anticipate its completion through the viewer.

For example, lately I’ve been thinking about pop art. I was in Pittsburgh and went to the Warhol museum. I was in Copenhagen and went to the Louisiana museum, where they happened to have a Marisol exhibition. Contemporary art has a special place in my heart, but I don’t like pop art. I never really have. However, I don’t think it’s fair to call it slop, even though it would fit many of the definitions above–it’s cheap, low-effort, could be produced in bulk, designed to mimic the predictable and forgettable. I can respect someone like Warhol because at the time, the work expressed a point of view, it contributed to a conversation, and by doing so opened a door to possibility, like all great art aspires to do. 

Slop, on the other hand, is talking when you have nothing to say. Slop is a waste of your time as a consumer, but also a waste of time for the author, who pleads for attention while denying themself a chance at discovering meaning. In this way, one could say slop is a matter of life or death, since after all, every moment is bringing us closer to death.

P.S. Merry Christmas to those who celebrate!