This is Jessica. There was a quiz on detecting AI-generated poetry making the rounds on social media yesterday that caught my attention. Poetry is a topic that I’m always curious about people’s perceptions of, since I suspect I have more knowledge of it than the average person. When I was younger I was very serious about it and even did an MFA in experimental poetics, at a program founded and directed by Beat poets Allen Ginsberg and Anne Waldman (though sadly Ginsberg was gone by the time I arrived in the mid aughts).
The quiz felt very easy – I got them all correct with little effort. Then I saw that it was created in response to a new Nature paper on how well people could detect AI-generated poems from a set (which included the poems in the quiz). They claim that “participants were more likely to judge AI-generated poems as human-authored than actual human-authored poems (χ2(2, N = 16,340) = 247.04, p < 0.0001).” I had figured I’d probably do better than the average person, but to think that the trend was actually reversed in the participants surprised me at first.
Here’s the abstract of the paper:
As AI-generated text continues to evolve, distinguishing it from human-authored content has become increasingly difficult. This study examined whether non-expert readers could reliably differentiate between AI-generated poems and those written by well-known human poets. We conducted two experiments with non-expert poetry readers and found that participants performed below chance levels in identifying AI-generated poems (46.6% accuracy, χ2(1, N = 16,340) = 75.13, p < 0.0001). Notably, participants were more likely to judge AI-generated poems as human-authored than actual human-authored poems (χ2(2, N = 16,340) = 247.04, p < 0.0001). We found that AI-generated poems were rated more favorably in qualities such as rhythm and beauty, and that this contributed to their mistaken identification as human-authored. Our findings suggest that participants employed shared yet flawed heuristics to differentiate AI from human poetry: the simplicity of AI-generated poems may be easier for non-experts to understand, leading them to prefer AI-generated poetry and misinterpret the complexity of human poems as incoherence generated by AI.
To anyone who reads or follows poetry, especially the more avant-garde kind, it’s well known that the kind of poetry that leads to accolades is often esoteric, challenging, and not necessarily all that pleasant to read. Part of the reason I lost interest in poetry myself was because at its peak it is so obscure, and the audience for it so small. I don’t think its a stretch to say that the point of what is considered “good” poetry to the experts is to subvert language, to defy expectations in a way that is confusing but opens up room for some unexpected sense of familiarity or recognition. This is true even of the classic, old-wealth New England scene that movements like the Beats were distancing themselves from. Consider Robert Lowell penning lines like “the glassy bowing and scraping of my will.” Or John Ashbery’s work, which put him at the pinnacle of established poetry, but also was a major influence on many outsider movements for his rejection of the conventions associated with establishment poetry.
In its more extreme forms, poetry forces the reader to “stop making sense” at all. Gertrude Stein was a master of this. For how long can one continue trying to apply the usual tools of reading to extract meaning when faced with something like this?
It is not a range of a mountain
Of average of a range of a average mountain
Nor can they of which of which of arrange
To have been not which they which
Can add a mountain to this.
Upper an add it then maintain
That if they were busy so to speak
Add it to and
It not only why they could not add ask
Or when just when more each other
There is no each other as they like
They add why then emerge an add in
It is of absolutely no importance how often they add it.
There’s a sense in which all poetry is deviant and for this reason, vulnerable even when it is most established or authoritative. I’m reminded of Anne Waldman frequently arguing that we must “Keep the world safe for poetry.”
So now I’m wondering, is poetry the antithesis to today’s large language models? Intentionally breaking typical structures of language so that the juxtaposition of each new line, or even word is somehow surprising to the reader’s expectations seems pretty opposed to what we expect to get when we combine autoregressive, predict-the-next-word objectives with post-hoc adjustment through procedures like RLHF, where people (not usually selected for their expertise) are shown many model outputs and asked to provide their preferences.
My colleague Matt Groh points out how results like the Nature study illustrate problems inherent to imitation game research: if you lack domain expertise and much knowledge of modern AI’s capabilities and limitations, it can be easy to fall for and even prefer simulacra. AI-generated content can be discernibly different from human-generated content, but in a way that seems more likable.
It makes me wonder how many other genres of writing there are where asking people how difficult or awkward writing is could be a good predictor for which is generated by human experts. Academic writing generated by ChatGPT is another example that seems fairly easy to detect to those with lots of domain knowledge, but where their preferences might be opposed to what non-experts prefer or associate with the genre.
P.S. This post reminds me of this paper we wrote on the role of aesthetic judgments in assessing the nature of modern AI, where we talk about how the ways we’ve learned to read art transfer to what we look for in AI outputs. It’s an interesting counterpoint to this discussion.
I’m not really sure what the useful conclusion of this paper is that it should be published in Nature. That the average person prefers “simpler” poetry that conforms to standards learned in the nursery? That AI produces this type of poetry more often because there is more of that in its training set (presumably lots of out-of-copyright Swinburne, Longfellow etc.?) That people pick AI poems because they “feel” like famous poems they know (and the AI training set includes famous poems they know including ones in copyright because they are often typed out in other places online)?
Also, I’m not sure that all good poetry breaks structures to the extent you claim: Larkin isn’t avant garde in structure or “next word unexpectedness”, but he is still
incredibly moving. Juxtaposition of ideas, images, and no doubt many other things beyond my ability to articulate are also part of poetry. I like the suggestion that poetry is the antithesis of LLMs, but I think that is for more reasons that you suggest.
I agree it’s not an earth shattering result. I was slightly surprised, but I don’t think my expectations about how easy it is to spot AI poetry are very predictive of the population. It does make me wonder though about the generating process that produces so much rhymed poetry when so much poetry does not rhyme. Is it mostly the training distribution, or is the fine-tuning using human preferences that happens later playing a role?
laypeople can’t distinguish AI-written from human-written poetry. yeah, obviously. there’s a reason why the judges on “is it cake?” are celebrities and not bakers.
It sounds like maybe they can though, if you believe their results (the detection rate was slightly below chance). They just get the label wrong. Probably lots of heterogeneity in the results though based on prior knowledge of what LLMs can do and what human poetry looks like.
I would love to see what kind of detection rate someone like Paul Hollywood has for Is it cake?
> is poetry the antithesis to today’s large language models?
As a non-expert in poetry, this strikes me as powerful reason to value the art in the world we currently inhabit.
Love this post. And at some point, I hope you share some of your poetry with me.
You reminded me of this great webpage from the aughties that generated emo adolescent poetry with next-token prediction.
https://www.elsewhere.org/journal/hbzpoetry/
I wonder how this would fare in the authors’ experimental gauntlet. The generation here is from a boring smoothed maximum likelihood language model. Details here:
https://www.jwz.org/dadadodo/index.html
The same blog also hosts a post-modernist text generator based on the same software. It captures the flavor of the worst excesses of 90s critical theory:
https://www.elsewhere.org/journal/pomo/
Yeah, I find this much more convincing as human-like poetry than what they used.
I remember the post-modernism generator. That kind of writing, plus the trendiness in experimental poetry, was what drove me back toward science.
To share my poetry I would have to find it first!
This all reminds me of how sometimes my submissions to scientific journals have been criticized as being too conversational, and how when the Monkey Cage blog moved to the Washington Post, our editor started telling me that my posts were too bloggy.
Style is important, and different outlets want different styles. Recently I read a nonfiction book by David Owen, who’s an excellent magazine article, but I found the book kind of annoying because it had various stylistic aspects—for example, describing each person’s physical appearance and giving some of their personal background—which is standard in magazine writing but which was often distracting to the main arguments in the book. Similarly, if I write in a pleasantly expository way, this can annoy journal reviewers who want to get straight to the core claim. And writing that is “bloggy” often has a kind of outsider feel, written as part of an ongoing conversation, which doesn’t work for newspaper editors who want to stand-along articles that can appeal to occasional readers.
Jessica’s point about poetry being challenging/different/oppositional is interesting, although I don’t know how true that is! From a historical standpoint, I see poetry as a continuation of oral traditions in storytelling, and also as songs without music. Popular storytelling tends to be accessible, not obscure. Songs are different: even popular song lyrics are often opaque and can benefit from decoding, which I guess is true of a lot of poetry too. But, to me, the obscure style in poetry is the norm. I have a horrible feeling that when people write poetry for the New Yorker or whatever, that they purposely take out words to make the poem harder to follow, and that the editors of the magazine are looking for that sort of obscurity. A readable, accessible poem in the New Yorker? Maybe, sometimes, just like sometimes there are scientific articles that are fun to read. But not usually, cos that’s not the rule.
That said, I guess that most people’s exposure to poetry is not from sources like the New Yorker or even classic poetry anthologies, but rather to greeting cards, inspirational poetry, and Bible verses, and all of those are in a much more accessible style. The subway sometimes has posters with poems, and they pretty much always seem to be in a weird halfway style: there will be a clear and uplifting message in the greeting-card style with some bit of obscure poetic formulation as in the New Yorker. I guess the idea is that, on one hand, you don’t want to annoy the subway riders; on the other hand, these are written by credentialed poets so they have to follow the rules.
So the problem of detecting “what is real poetry” is gonna be tricky, given all the constraints of style.
I was thinking about greeting card expressions as a better analogy for what an LLM calls poetry than poetry.
On taking out words to make it harder to understand… apparentlly W.H. Auden, who was the judge for a prize that John Ashbery won for his first book, admitted later that he didn’t really understand any of it. Ashbery continued winning prestigious awards after that, including a Pulitzer. So who can blame the New Yorker poets!
I agree “real poetry” is hard to define. But “poetry as recognized by people society would consider poetry experts” is less ambiguous.
Lots of modern art is reflective within its own genre, and simultaneously looking for novelty, often on a meta level. LLMs don’t have much training material of these “language games”. There’s that tendency to break predictability as you say, to deliver novelty by breaking old representations while still being meaningful. It may indeed be antithetical to LLMs. Also in that embodiment plays a role: part of that breaking presumably comes from non-linguistic patterns of physiology (feelings), psychology, relationships, environment.
About the sample poem, an LLM says: “I don’t get the sample poem either, despite my earlier attempt to find meaning in it.”
Gertrude Stein would probably be pleased to hear this feedback from the LLM
Humor is another type of writing that’s often held out as something AI is going to find hard to crack. It’s not necessarily “difficult or awkward” to read in the way that some poetry or academic writing can be, but (acknowledging there are people spending a lot of time trying to define exactly what makes something funny) a fair amount of humor gets there by taking a narrative in unexpected directions.
“Nor can they of which of which of arrange”
This reminds me of DFW, who suggested a fun exercise for precocious youngsters to come up with a valid sentence that used the word that, five times in a row. His solution: “He said that that that that that journalist used, should have been a which”
Love this.
In linguistics and philosophy, my old field, we’d score that as four uses of “that” and one mention (hint: the third “that” should have been quoted). See: https://en.wikipedia.org/wiki/Use–mention_distinction
The author was right on the standard “which” vs. “that” convention.
I’m sure you could reproduce this experiment with art. That’s also something that’s become so meta and self referential that it’s more about what’s written on the wall than what’s actually being shown. No way I could tell the real stuff from a made-up spoof other than if the shelf-talker didn’t scan as convincing art-speak.
This also reminds me of taste tests done on wine with groups of novices and experts. The novices can’t guess which wines are more expensive, but the experts can. There have been a bunch of studies where this pattern emerges—there was quite a large one about ten years ago that was going around then and most people just cited the topline (people can’t tell expensive wine from cheap wine).
It also reminds me of Andrew’s continued warnings about interpreting group averages.
Ah yes I hadn’t thought of wine but that’s a great example.
Agree, the group average is probably hiding a lot here.
Bob –
There have been a bunch of studies where this pattern emerges—there was quite a large one about ten years ago that was going around then and most people just cited the topline (people can’t tell expensive wine from cheap wine).
Do you have a link by any chance? My memory was trials showed explicitly that wine experts didn’t identify the expensive wine. It did seem very counterintuitive and I thought it showed a provocative generalizeable pattern, and I’d like to know how I got that wrong. As I recall, there were some trials that supposedly shows experts couldn’t tell red wine from white wine which was REALLY hard to believe.
I was about to give up, then I asked ChatGPT 4o to find what I was looking for. Here’s a nice meta-analysis:
https://wine-economics.org/wp-content/uploads/2012/10/Vol.3-No.1-2008-Evidence-from-a-Large-Sample-of-Blind-Tastings.pdf
Excerpts (my emphasis): “A number of studies have reported positive correlations between price and subjective appreciation of a wine for wine experts (e.g., Oczkowski, 1994; Landon and Smith, 1997; Benjamin and Podolny, 1999; Schamel and Anderson, 2003; Lecocq and Visser, 2006). Non-experts, however, may not be particularly sensitive to some of the refinements that are held in high esteem by wine aficionados.” And “Our main finding is that individuals who are unaware of the price do not, on average, derive more enjoyment from more expensive wine. In fact, unless they are experts, they enjoy more expensive wines slightly less.”
These actually seem to contradict each other a bit. If they enjoy expensive wines slightly less, then they could use their preferences to detect the difference.
A similar blog post on ACX, but for images instead of poetry: https://www.astralcodexten.com/p/how-did-you-do-on-the-ai-art-turing
The “non-experts” bit gives the game away: poetry, jazz, art, classical/formal/serious music are things that you need experience with to understand. (That is, the language in which an art form speaks has history and vocabulary and rules that must be learned.)
Also, although not mentioned in this post, another discussion of this article mentioned that a lot of the human-written poems were older ones by the usual dead white male suspects (oops: major authors), and that’s writing that current folks have less familiarity with than folks of my generation.
Personally, I’d guess I’d do badly on such a test. I find rhyme and meter to be irritating and trite, so don’t “do” poetry. Haiku as done in Japan, though, is pretty neat. Despite the rules being rather formal, said rules actually do help in creating compact works that have a lot to say for themselves.
Interesting quiz. Thanks for posting about it!
For me, what made the LLM generated poems in the quiz seem “wrong” wasn’t their lack of linguistic inventiveness, though — they were just bad poems. They often scanned badly, or had forced seeming rhymes or dull imagery. On a couple of them the diction just seemed off — the fake Whitman and the fake Plath in particular had a sort of naive self-centeredness to them which felt like the LLM had been trained on some teenage writing.
It may well be true that a large portion of poetry available on the internet for LLM training has been written by adolescents.
I got an 8/10. I marked the most subversive one actually as AI (thinking it was intentionally trying to write a confusing poem). I don’t know anything about poetry, but for the most part, I just marked poems as AI if the beat didn’t feel right when I said the poem aloud to myself
The Nature Scientific Reports paper (note: not Nature proper!) lacks perspective: there’s a history of studies testing people’s ability to understand and evaluate poetry, a discussion of which would have provided helpful context on the extent to which randomly selected survey participants might be expected to be able to read Chaucer or Dickinson. The study also had some issues, which I noted in this PubPeer thread.
Thanks for pointing out that its not Nature proper, and sharing the critique.
I talked about this study with a friend of mine; the results are apparently shocking, but I think they make sense. 1) most people do not read poetry, they find it confusing, hard etc. Poetry reading requires training. It is hard actually 2) They way LLM works is by maximizing likelihood, so it will always go (give or take some noise) for a construction that has high frequency. This is why, when it is asked to give some input on fiction or poetry, it falls on cliches (I’ve picked the explanation from Nassim Taleb; at some point someone asked Chat GPT to provide the plot for a novel, and the engine suggested several plots, none of which were right, because all of them relied on common places). 3) I’ve also played with ChatGpt, prompting it to write poetry. I tried several poets, and I was disappointed each time by the bland output. 4) I think some of the older, experimental poetry will look even worse in the AI era. I think Chat GPT would be able to quickly generate a version of the Exercises in style by Queneau, really similar to the original. It is funny because in the past I was interested in Oulipo techniques, I did collages like Raymond Queneau, Tzara and Ted Berrigan (but being a statistician, I used the R command sample to shuffle sonnet lines).
Yes, I agree they make sense. Poetry reading is hard.
I remember Oulipo, and that there were a few automated (rule-based) poetry generators back when I did my MFA in poetry. One of the first things I did when I first learned Java was create a haiku generator.
This post reminded me of a discussion I had with someone from a song-lyric forum who seemed a big fan of using AI for lyric writing. I am not a fan of that, and I tried to make clear that I don’t care that much if my writing uses simple rhyming or whatever AI can possibly improve upon, but that I care about evoking some emotion in some way.
One of the things I wonder concerning AI is whether it can produce writing in which the interpretation of stuff is not clearly tied to the words. This can be done in several ways I think, and I hope it’s okay to share a poem I wrote in which the interpretation is not totally clear from the use of the words. It sort of mixes up birds and a scenario involving two people. I still don’t know if the poem makes sense, or if the reader understands what it might be about. It may be a nice example of what I am trying to make clear concerning AI writing, so I hope it’s okay for me to share the following:
Two Birds
I flew away, or I let you fly away
I still don’t know exactly what to say
Maybe your song was too beautiful for me
Or maybe we sort of set each other free
I thought of your chirps over time
It sometimes seemed like a song without rhyme
But maybe your chirps in their entirety
Constituted a song with a beautiful melody
Then all of a sudden on one particular day
I heard your call in a place not far away
I answered back with a particular sound
At that moment, it was the best I found
Apparently it wasn’t easy chirping along
After hearing your unexpected song
Once more I wonder whether it could be
That it might just be too beautiful for me
Still, I hope to hear your chirp again one day
But, if you will, you have to chirp from far away
If I do not answer back immediately
I am looking for a better spot to hear and see
In the meantime I am cleaning my feathers
And searching for a place that better weathers
If I fly now I will likely fall to the ground
I want to chirp back when I can make a better sound
So, if you want to, you could someday chirp again
And I will likely answer back when I can
Then, if you will, you can hop to the next tree
And I’ll do the same, to better hear and see
Then on one morning when you are near
I will sing a song precise and clear
You will know it is meant for you
And I will know it is the best I can do
Quote from above: “One of the things I wonder concerning AI is whether it can produce writing in which the interpretation of stuff is not clearly tied to the words. This can be done in several ways I think,(…)”
I am the same Anonymous that posted the above comment which included the poem titled “Two Birds”. I had a discussion about poetry and AI a moment ago with someone. I was trying to make clear that I wondered whether human intelligence can create poetry or song lyrics that AI can’t (I think a slightly better way of phrasing might be wanted here but I can’t think of the words, and I hope you get what I mean).
I used the example of having written a poem/song lyrics in which two words “I know” in the verses have different meanings and interpretations (and possibly even different imagined pronunciations) which relate to what’s being depicted in the verse. I wondered whether AI can produce such writing, and remembered some discussions about such things on here. I did some searching and found this blog post which seemed an appropriate place to try and get some answers concerning my wonderment.
It’s another example, just like the content of “Two Birds”, where I think human writing and human interpretation is necessary for the poem to “work” and I sincerely doubt AI is capable of writing such things. I however know little of AI’s possibilities regarding this, so I thought I’d see whether anyone has some answers for me concerning this all.