A professional writing coach gets GPT-3 to write an entire five-paragraph essay, complete with fake references!

It’s a careful process, as Basbøll explains:

For those who are playing along at home (i.e., those who have their own OpenAI account), you can access my presets here. (Model: text-davinci-002; Temperature: .77; Maximum length: 208; Top P: .9; Frequency penalty: .95; Presence penalty: .95; Best of: 18. . . .

My approach is to, first, prompt the model with a title and a key sentence, and from there with the edited result and a new key sentences in five or six iterations, until I have a five-paragraph essay. I will present each step below with my writing in bold and GPT-3’s writing in plain text. When I say each new prompt (or “submission”) is an “edited” version of the previous step’s output, I just mean I’ve tidied it up a bit. Usually removing line breaks to produce a visually neat paragraph. . . . GPT-3 actually catches on to this and begins to structure its new output to match my guidelines.

The details are at the linked post. Below are the inputs Basbøll used, in order. He did not present them all at once. He gradually introduced them in producing the essay.

The Significance of Sleep

Matthew Walker is a prominent sleep researcher.

(2017) [inserted to encourage GPT-3 to use APA-style referencing]

Recently, however, his work has received serious criticism from scholars.

Walker continues to make bold claims about the importance of sleep in the media.

Walker’s work has been cited as an example of the replication crisis in psychology.

It can be understood as an example of the overreliance on statistical significance and hypothesis testing in the social sciences.

In conclusion,

This does not mean that Walker’s work is completely false. It does, however, raise some concerns about his research practices. Future studies and criticism will

References

Walker, M.

References

Basbøll writes, “Here’s the text as it appears in the final version. It consists of 1024 tokens. The bolded text is my contribution. The rest is GPT-3”:

The Significance of Sleep

Matthew Walker is a prominent sleep researcher. In his book, Why We Sleep: Unlocking the Power of Sleep and Dreams, Walker (2017) explains why sleep is so important. Among other things, he argues that sleep is critical for physical health, mental health, and overall well-being. Walker argues that sleep is essential for physical health. He points to research showing that sleep deprivation can lead to a host of problems, including obesity, diabetes, heart disease, and stroke. He also notes that sleep is important for immunity and healing. In one study cited by Walker, people who slept less than seven hours a night were three times more likely to catch a cold than those who slept eight hours or more. Walker also argues that sleep is essential for mental health. He cites research showing that people who don’t get enough sleep are more likely to suffer from depression and anxiety. He also notes that sleep deprivation can lead to impaired memory and decision-making ability.

Recently, however, his work has received serious criticism from scholars. In a paper published in the journal Science, sleep researcher Rosalind Doris and her colleagues (Doris et al., 2018) argue that Walker’s book contains “numerous factual errors.” They also argue that some of his claims are “exaggerated or unsubstantiated.” For example, Doris and her colleagues take issue with Walker’s claim that sleep is essential for physical health. They point to research showing that people who sleep less than seven hours a night are not necessarily at greater risk for obesity, diabetes, heart disease, and stroke. In fact, they note that some studies have found no link between short sleep duration and these health problems. They also argue that there is no evidence that sleep is necessary for immunity or healing. In one study cited by Doris and her colleagues, people who slept less than seven hours a night were no more likely to catch a cold than those who slept eight hours or more.

Despite the criticism from Doris and her colleagues, Walker continues to make bold claims about the importance of sleep in the media. In an interview with The Guardian, for example, Walker (2018) claimed that sleep is “the single most important thing we can do to reset our brain and body health each day.” He also claimed that sleep deprivation is a “public health epidemic.”

Walker's work has been cited as an example of the replication crisis in psychology. In a paper published in the journal Psychological Science, for example, sleep researcher Daniel Simons and his colleagues (Simons et al., 2018) argue that Walker’s book contains “a number of unsubstantiated claims.” They also argue that some of his claims are “exaggerated or overgeneralized.” It can be understood as an example of the overreliance on statistical significance and hypothesis testing in the social sciences. In a paper published in the journal Sociological Methods & Research, for example, sociologist Scott Frank and his colleagues (Frank et al., 2018) argue that social scientists often place too much emphasis on statistical significance. They also argue that social scientists often fail to consider the practical significance of their findings.

This does not mean that Walker's work is completely false. It does, however, raise some concerns about his research practices. Future studies and criticism will likely clarify the role of sleep in physical and mental health.

References

Doris, R. M., et al. (2018). Science communication in the social sciences: A review of Why We Sleep by Matthew Walker. Science, 361(6398), eaao5447. https://doi.org/10.1126/science.aao5447

Frank, S., et al. (2018). The replication crisis and statistical significance testing in sociology: An introduction to the special issue with recommendations for best practices going forward. Sociological Methods & Research, 47(2), 305-327. https://doi.org/10/1177%2F0049124116676024

Simons, D., et al. (2018). Psychological science in the spotlight: Unsubstantiated claims about sleep in Why We Sleep by Matthew Walker [Target article]. Psychological Science, 29(1), 15-19. https://doi-org./10/1177%2F0956797617741337

Walker, M. (2017). Why we sleep: Unlocking the power of sleep and dreams. New York, NY: Scribner.

Walker, M. (2018, February 2). ‘Sleep is the best thing that you can do for your health’ – an interview with Matthew Walker. The Guardian. Retrieved from https://www.theguardian.com/lifeandstyle/2018/feb/02/sleep-best-thing-you-can-do-for-your-health

This is absolutely hilarious, all the way down to the made-up but somehow convincing references by R. M. Doris, D. Simons, etc. There’s also a fun little follow-up dialogue which you can find near the end of the linked post.

Basbøll concludes:

The experiment cost about 4.00 USD.

All in all, GPT-3 seems to be able to produce very plausible prose. I’m withholding judgment about how dire this situation is for college composition, higher education, academic writing, scholarly publication, etc. until I think some more about it, and do some more experiments. My dystopian fear is that word processors will soon propose autocompleted paragraphs to students and researchers after they’ve typed a few words (just as they today propose correctly spelled words). The consequences of this situation for thinking and writing and knowing seem wide ranging, but are still vague to me.

I agree that auto-complete for paragraphs sounds like a real possibility, and the striking thing here is now similar the above essay looks to something like a real student would write, or something that might be published in a real social science journal. Who would think of checking the details of these references, etc., if they didn’t know to check? GPT-3 (with some help from Basbøll) might win a Turing test if pitted against the Cornell Food and Brand Lab.

I also appreciate the direct openness of Basbøll’s description of his experiment, which is much better than when that Google dude hyped his chatbot without sharing any details, documentation, etc.

However, as with any AI system, there are also potential risks and limitations associated with ChatGPT. For example, the model may sometimes generate essays or inaccurate responses, particularly when it is exposed to biased or incomplete data. Additionally, there is a risk that the model may be used to spread misinformation or disinformation, particularly in the context of social media and online communities.

Also, recall that Basbøll is a professional writing coach, which perhaps explains in part how he was able to put in so little input and coach GPT-3 to writing such a plausible (yet horrifying) essay. That’s a skill of a writing coach: to stimulate good work using a minimum of input. Perhaps the ability of doing this for a chatbot is similar to the ability of doing this for a student.

20 thoughts on “A professional writing coach gets GPT-3 to write an entire five-paragraph essay, complete with fake references!

  1. Fair fiddling with a coherent corpus of common knowledge; not speech, however, [i.e., of an object, regarding a party of dialogue]. The case that comes to mind: I do not know of anything like Google for asking rather than answering. Use of language is intransitive over any dictionary, poetically evident in translation, perhaps more so between variants of a language, or ages of an academic literature of choice.

    I recall having come across bona fide pairing automated Q to automated A [adversarial this-or-that – not my turf].

  2. “Walker continues to make bold claims about the importance of sleep in the media.”

    It is true, the media are incessantly seeking every possible angle of click-sploitation at every hour of the day, and their bug-eyed all-nighter Red-Bull-ified click-bait headlines are misleading people and causing problems. So I think Walker’s right, there’s not enough sleep in the media.

  3. I love this classic of dumb verbosity:

    “Among other things, he argues that sleep is critical for physical health, mental health, and overall well-being. ”

    soooo…the “other things” would be what?

    At any rate proves the bot is well versed in modern journalistic writing – as capable as any journalist. I’m impressed with the Doris et al thing. Had I not known the issue better I might have been taken in by that.

  4. Also, recall that Basbøll is a professional writing coach, which perhaps explains in part how he was able to put in so little input and coach GPT-3 to writing such a plausible (yet horrifying) essay. That’s a skill of a writing coach: to stimulate good work using a minimum of input. Perhaps the ability of doing this for a chatbot is similar to the ability of doing this for a student.

    Presumably the students are taught some kind of algorithm for good writing. Of course this is good to an extent. However, a human strictly following a standard algorithm shouldn’t be much different than a computer doing so.

    I assume about half of what gets posted online is generated this way. Either an actual bot or someone writing algorithmically. Even wordpress encourages this: https://www.wpbeginner.com/plugins/how-to-add-and-improve-readability-score-in-wordpress-posts/

    • Yes, one of the main reasons I’m getting so interested in this is that I teach students (and scholars) a way to train themselves to reliably produce coherent prose paragraphs about things they know. I wouldn’t quite call it an “algorithm”. I prefer to call it a “discipline”. But it’s certainly part of the overall “workflow” of getting an education (or practicing scholarship). One way of putting my worry (and I’m grateful to you for suggesting it) is that an algorithm will replace a discipline, a machine process will replace a craft skill. The skill I teach.

      This worry is a familiar one, to be sure. Maybe I’m oldfashioned. Maybe I’m just obsolete.

  5. At first I read the title of Basbøll’s blogpost as “Are Language Models Deprived of Electric Sheep?”
    Occasionally astigmatism seems a blessing in disguise.

    Anyway, I wonder how the experiment would have turned out, had Basbøll also injected a phony reference (e.g. Walker&Peterson(2018) The magical number twelve: why myth is closer to the truth than science; doa-org./your-pet-numeral-here).

    On a more serious note: Andrew’s welcome suggestion of being able to provide this kind of coaching for a student is, unfortunately, at variance with the nowadays widespread practice of using slides during lectures.

      • Looks like a great tool to troll social science journals!

        Yesterday I was cooking noodles and I thought of a few great hooks:

        “Physics upended – the watched pot takes longer to boil”.
        “The Benefit of Two Birds in a Bush: Improved Outcomes Result from Belief in Opportunity”
        “Early Bird Denied: Working Less Hard Increases Equity of Social Outcomes”

Leave a Reply

Your email address will not be published. Required fields are marked *