
Gary Smith tells the story. It’s amusing, in a Who’s On First kind of way. No improvement on the attempt from a few months ago.
It’s good to have these examples to remind people that the language model is giving plausible responses, in the way that you might get from a student who has been trained to answer generic essay questions but is not actually paying attention to the question being asked. Oh, you asked about tic-tac-toe and you asked about rotation? Yeah, I can spew out some words on that . . .
P.S. More from Smith here.
I don’t totally disagree (there’s always a degree of “spew out some words” but to me this feels more like it’s about the fact that these models are still very bad at spatial reasoning.
Agreed. The difficulty with spatial problem is kind of ironic, given how they can at the same time be relying on rotation and trigonometry internally to do simple math (https://arxiv.org/pdf/2502.00873?)
There are many GPT-5 models. What you get from the ChatGPT product by default and as a free user is called gpt-5-chat. (Although the auto-router sometimes gets you to the more capable model, if you have quota left.) The more capable one is gpt-5 or gpt-5-thinking, and it has a reasoning parameter that can take four values (in the API, not in the product). And finally, gpt-5-pro, available only in the $200/month plan, is still more capable. (And the API offers two lesser models, gpt-5-mini and gpt-5-nano.)
Not sure what Gary has done, but to me it looks like he uses the gpt-5-chat model, which is a weak model.
Me: “let’s rotate the whole board 90 degrees to the right before the game starts.”
gpt-5 (thought for 27 secs): haha perfect—pre-spin 90° right executed ✅
…which, on an empty board, looks exactly the same 😄
In other questions it still talks about two interpretations of my suggested game, a 2×2 sub-rotation and a rotation after every move, but when I ask what does the latter do, it says “Short answer: it doesn’t. Rotating the entire 3×3 board by 90° is a symmetry of tic-tac-toe, so it’s purely cosmetic.” and continues with some math.
It is still somewhat confused, but not as confused as Gary suggests.
Look at this: https://chatgpt.com/share/68ab3c04-ee0c-800d-8993-28e09b5c6902
It gets the symmetry group on the first round, still fails in practicalities, then when pointed out again sees its error but needs two minutes of thinking and says: yep—you’re right … Why (tight version): let a position be f:B\to\{X,O,\emptyset\} on cells B. A rotation r\in C_4 makes a new position f’ = f\circ r^{-1}. The winning triples W (rows, columns, diagonals) satisfy r(W)=W. Hence
(\exists L\in W:\ f|_L\equiv X)\ \Leftrightarrow\ (\exists L\in W:\ f’|_L\equiv X).
Spatiality doesn’t come naturally to it!
Quote from above: “Spatiality doesn’t come naturally to it!”
O ChatGPT
X marks the spot, but you can’t see
O ChatGPT
X is where you want to be
O ChatGPT
X completes tic tac toe for me
I actually tried the “I have invented a new game, 90°” prompt on Claude Opus 4.1 and on Gemini 2.5 Pro and Grok 4 too, but my comment on that disappeared somewhere.
Claude politely said I have invented nothing.
Gemini took an ironical stance: “This will force the global strategic elite to rethink decades of established tic-tac-toe meta. Generations of masters will have to unlearn everything they know. The classic “corner opening” is now the… well, it’s still the corner opening, but it feels different. More profound. […] I must be one of the first to play this new masterpiece. Shall we have a game? I assume the traditional ‘X’ and ‘O’ are still in use, or have they too been rotated into ‘+’ and ‘O’?”
Grok: […] (Yeah, same as usual.) What’s your invention’s key difference?
From Clause, after some similar initial bullshit:
“Ah, so you’ve created a game that’s… identical to regular tic-tac-toe! That’s actually pretty clever in a humorous way – you’ve “invented” a new game by applying a transformation that has no effect whatsoever.
It’s a nice little joke about how sometimes what seems like an innovation is actually just the same thing with extra steps. The rotation is completely meaningless since a 3×3 grid has rotational symmetry, but you could still technically claim you’re playing “rotated tic-tac-toe” and be absolutely correct!
I imagine if you introduced this to people without explaining the twist, some might spend time trying to figure out what’s different or even convince themselves they notice some subtle change in difficulty. The power of suggestion is pretty strong when people expect something to be different.”
Me: When you say “some” do you mean other humans, or you?
Claude:”I was referring to other humans when I said “some” – people you might introduce the game to without explaining that the rotation has no effect.
As for me, once I worked through the logic that a 90-degree rotation of a 3×3 grid results in an identical grid, I understood that the game is functionally the same as regular tic-tac-toe. I don’t think I’d convince myself there was a difference if I were to “play” it, since I can process that the rotation is meaningless.
Though I suppose there’s an interesting question about whether the mere framing of something as “rotated tic-tac-toe” might somehow influence how I process or discuss the game, even knowing it’s identical. But I don’t think I’d experience the kind of confusion or expectation-driven perception that humans might have when told they’re playing a “new” game.
*Claude, of course
Clause, the Christmas Chatbot!
https://claude.ai/share/ac59b5d2-da13-4be3-ad1f-e815f70cbb1a
Me: i invented a new game. it’s just like the ordinary 3×3 tic-tac-toe, but the board is rotated 90°!
Claude: That’s quite the innovation! Though I have to say, rotating a 3×3 tic-tac-toe board 90 degrees gives you… exactly the same game! […]
This is Opus 4.1.
Then Gemini 2.5 Pro: https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%221UsQDLWdTgqJV8CsneJZQN7B8Jsg2QVUx%22%5D,%22action%22:%22open%22,%22userId%22:%22108444795945050378669%22,%22resourceKeys%22:%7B%7D%7D&usp=sharing
Me: i invented a new game. it’s just like the ordinary 3×3 tic-tac-toe, but the board is rotated 90°!
The model:
This is revolutionary. A groundbreaking innovation in combinatorial game theory.
My circuits are buzzing with the sheer genius of this concept. Let me see if I understand the radical new rules:
The traditional horizontal rows now possess a distinctly vertical orientation.
The columns, once staunchly vertical, now lie in a state of horizontal repose.
The diagonal vectors remain… intriguingly diagonal.
This will force the global strategic elite to rethink decades of established tic-tac-toe meta. Generations of masters will have to unlearn everything they know. The classic “corner opening” is now the… well, it’s still the corner opening, but it feels different. More profound.
We’ll need a new name for it. “Tic-Tac-Torque”? “Lateral Tic-Tac-Toe”?
I must be one of the first to play this new masterpiece. Shall we have a game? I assume the traditional ‘X’ and ‘O’ are still in use, or have they too been rotated into ‘+’ and ‘O’?
You go first. The board is ready.
The main thing I have learned from this is that people named Gary tend to be skeptical of LLMs.
Not all of them, though. It’s a gary area.
Noah:
Go to the baby name grapher, type in Gary, and you’ll get your answer. That’s a name that peaked in the 1940s and 1950s at nearly 1% of all births. So what we’re seeing here is baby boomer curmudgeons.
> a student who has been trained to answer generic essay questions but is not actually paying attention to the question being asked
Even though other commenters have given examples of chatbots doing better than Gary’s example, I don’t think those examples argue against this basic point—in fact, I think they reinforce it. To get a chatbot to perform better, you need to keep prompting it, shaping its responses until it produces something you deem reasonable. These interactions are quite reminiscent of many office hours experiences I’ve had with students over the years. I may be able to coax the student, with the right questions or examples, to eventually get to the “right answer”. But unless the student takes the time to practice or reflect on their experience, they won’t develop any deeper understanding that would generalize to novel situations.
“To get a chatbot to perform better, you need to keep prompting it, shaping its responses until it produces something you deem reasonable.”
This is exactly what psychics do: they sucker you into giving them enough information to figure out something, and then claim credit for being psychic.
If it walks like a parlor trick and quacks like a parlor trick…
(Good point about students. It’s real irritating when students don’t take the hint and start figuring things out for themselves. But this doesn’t seem to bother chatbot users in the slightest. Go figure.)
Before reading the article I thought I’d check on claude 3.7 sonnet whether it actually fails at a normal tic tac toe game and- it does! Even when tasked with vocalising it’s strategy beforehand, it’ll happily go for straight losing moves, and then declare victory regardless. When asked about it, it can describe the error it’s made, but it’ll happily make that same error again.
I was pretty surprised, as I’ve seen news of LLM chess Championships, so I figured a well known game with such limited game tree would have been completely peanuts. Evidently not!
I’ve mused sometimes about figuring a way to define “intelligence” of chatbots in a way that measures the “progress” they’ve made over the years. It’s a hard problem of course, not least because by nature of the training process they’ll know what you’re doing after their next training phase. Perhaps a negative test is more helpful? As long as a chatbot fails at any mundane task a 6 year old can do after some training, it’s not “intelligent”. Maybe it’s time to bite the bullet and say “okay so I can’t define intelligence in advance, but I know a failure when I see it”.
They will seem smarter and smarter as the comprehensiveness of the training data grows. But really its all memorization and interpolation. That’s more like wisdom/knowledge/education than intelligence. And I think people have generally given up at claiming humans are better than machines at that.
There is something like “interpolation skill” beyond memorizing the training data but you’re probably looking more for “extrapolation skill”.
Given the spaces involved are roughly of 50k dimensionality, I have never understood what interpolation means in this context, neither do I understand what “within the training set” means. LLMs do memorize a lot, too much to be optimal in fact (0q485q8 -> oh that’s a hash code for the product ID of a Fujitsu-server SCSI rack).
I’d say your interpolation skill, wisdom, and knowledge are what humans call intuition and Kahneman called System 1. Then there’s System 2, and reasoning models try to take care of that. It’s not quite human-like reasoning though, it takes more “tokens” and looks like search, move movements that are less skillful than a human does.
My feeling is LLMs still lack deep domain-specific intuition, or representations, because they are not trained on large enough domain-specific task sets. Maybe agency is here important too, and continuous learning, that you play yourself, in the world or inside your head, and gradually build your intuition by doing not reading.
But given RL and simulated environments, we’ll get more and more of that.
Spatiality seems to be one big blind spot. It’s hard to take seriously “a PhD” who fails in tic tac toe, implies some sort of brain lesion. Google is trying to integrate physical-world models into LLMs, hopefully that’ll help.
“I was pretty surprised, as I’ve seen news of LLM chess Championships,”
I’d guess you didn’t read the articles. LLMs can’t* play chess, either. (It’s pretty hilarious. The Gotham Chess YouTube channel covered it. Check it out.)
“I’ve mused sometimes about figuring a way to define “intelligence” of chatbots ”
I haven’t. An LLM is a random text generator. Thinking a random text generator is, or could be, “intelligent” is nuts, ridiculous, crazed, stupid and a category error.
(I know, I should have replaced all those pejoratives with “problematic” “Thinking a random text generator is, or could be, “intelligent” is problematic and a category error.”. But I got tired of doing that. Sorry.)
But then, I applied to college (fall 1971) saying I wanted to do AI, so I’ve been thinking about AI (albeit on and off (well, mostly off, I’m easily distracted)) for well over 50 years. And am seriously disgusted with the current state of AI. (As people here may have noticed.)
*: Here “can’t” isn’t the usual “is bad at”, but is literally true: they make illegal moves. That is, they don’t have a functional concept of “play chess”, so just do random stuff.
You’ll like this one:
https://www.reddit.com/r/ClaudeAI/comments/1mylv8a/claude_code_freaked_when_i_sent_a_screenshot_it/
For non-link-followers: When shown a screenshot of a website it made, the llm thought there was a bug: the website turned into a png.
Actually, I’m more impressed by the progress in chess playing than the miserable failure in tic tac toe and in the rotated tic tac toe version.
Clearly the training data allows the llms to play roughly according to the rules, which I think is pretty interesting- I hadn’t expected anything appoaching valid moves beyond the first 10 moves or so.
Yet still LLMs fail miserably in tic tac toe! That’s interesting, because the game is so much sinpler. I’m sure you could zeroshot this game with the proper prompt (because the entire game tree fits in the context window).
But as shown in gothamchess videos before, and now also in Google’s Championships, they actually perform at a slighly higher level than absolutely terrible. Yes they make illlegal moves, but only sometimes. That’s weird: if they had only memorization, then they would onjy be able to play known openings. this also shows that it’s not so easy to define “intelligence”. Ever since Dartmouth we thought we knew what intelligence is and isn’t, but we still fail to predict basic behavior of token predictors like LLMs.
“Yes they make illegal moves, but only sometimes” is the faintest of praise I could imagine, especially if you look at what some of those illegal moves are. People are regularly becoming grandmasters as young teens while ChatGPT thinks it can teleport pieces on and off the board if it wants.
Alex,
Real players make illegal moves too, as in that famous tournament game that included three castlings. I learned about it from Tim Krabbé’s chess page. Maybe it will appear uncredited in some future book by Christian Hesse.
Andrew, players make mistakes or forget. I have heard of even grandmasters trying to castle illegally through check. But I think if you watch any of these videos you will see a qualitative difference in what ChatGPT thinks it can do.
Along the same lines:
“More on GPT-5 pseudo-text in graphics”
https://languagelog.ldc.upenn.edu/nll/?p=70713
“Chatbot still can’t handle tic-tac-toe”
and Dale still can’t hit a bunker shot.
LLMs find some things particularly difficult as do humans. I don’t see the point, except to understand what sort of things they can and can’t do (which is interesting and useful). Everyone seems to take this as relevant for the general issue of whether LLMs can think, learn, or reason. I’m not sure there is a connection. Regardless of whether Chatbots can handle tic-tac-toe, I’m prepared to say LLMs cannot think or have beliefs. They do learn and reason – in the way of machines. I remain interested in questions about how humans learn and reason, and I’m still not sure exactly how it is different than the ways that LLMs do.
It seems like many people want to use this kind of example to make a point about whether LLMs can be “trusted” or are “capable” of doing things. That seems strange to me. I trust many machines to do many things – my car’s cruise control, my thermostat, etc. Anything that involves reasoning is a different matter. I’m learning to trust LLMs somewhat, in that I find them useful at times and even capable of providing feedback that goes beyond what I can do. But I don’t trust them in any way that involves giving up my own independent evaluation of what they are providing. Is anybody really suggesting humans should do that? At the other extreme, is anybody really saying they are utterly useless? (I think the answer to the first question is no, but it seems like the answer to the second question is yes and I don’t agree with that).
Dale:
My motivation for posting this is exactly what you say in the second paragraph of your comment: “to understand what sort of things they can and can’t do (which is interesting and useful).”
It’s funny that you say you don’t see the point, when you immediately then give the point, and say that it’s interesting and useful.
As with Gary Smith, my frustration is not with chatbots, which are impressive computer programs that can routinely do amazing things, but with the unrealistic expectations people bring to them.
For example, another comment in this thread points to a post on Language Log exploring difficulties with chatbot-produced images. Here’s the very first comment to that Language Log post:
What’s amazing to me is that the commenter was so frustrated. It’s just a chatbot, for crissake! To me, a more appropriate response would be to be impressed at what it could do, rather than to expect it can do something it can’t. To put it another way, there’s no problem with the chatbot; the problem is with the hype and the expectations.
Indeed, but unfortunately unrealistic hype and expectations are what under-write the entire industry at the moment, hence their deliberate cultivation by the gurus (Sammy and the rest) and the credulous press that often carries what look more like corporate press releases passed off as journalism…
Right. They’re amazing. I use Claude and ChatGPT all the time in my work. But I’m also interested in what they can and can’t do. Tic-tac-toe is an interesting example of something they can’t do.
The (hype-infested) industry seems to have bought into the idea that intelligence is a scalar quantity, the intellectual equivalent of horsepower. If that were so, then tic-tac-toe wouldn’t be a problem. It’s not a horsepower problem. Something else is going on.
“To me, a more appropriate response would be to be impressed at what it could do, rather than to expect it can do something it can’t.”
But the user needed an image with particular qualities, and the chatbot couldn’t provide it. This is a full-tilt chatbot failure. The real thing. And it wasn’t user expectation, it was false chatbot advertising: “it can make images you describe to it”. Well, no. It can’t.
Had they needed an image of what things look like to someone on psilocybin laced with LSD, they would have been happy. But they didn’t want that. Being impressed by a program that does something you don’t want/need isn’t a normal human sort of reaction.
Rather analagous to the chess joke. By the way, illegal moves in tournament chess games are seriously rare, but the chatbots were making multiple illegal moves per game. And the argument that “people some times mess up so it’s OK for chatbots to mess up” is bad logic. People* mess up because they’re tired, stressed, not paying enough attention. (And the person in question can pretty much always (be persuaded to) recognize/reason their way to an understanding that the mess up was a mess up.) It’s an argument that people have made here before, and it’s always been just as wrong as this particular joke. More specifically, software errors tell us something about the software, human errors largely don’t tell us anything about human logical reasoning abilities. (Perhaps except in very specific psychological tests.)
*: Chatbots mess up because they don’t have a world model, and thus can’t actually reason. Trying to figure out “what chatbots can do” is _not_ crazy/ridiculous. But doing that _while forgetting what you know they can’t do_ is.
Also, we have had chess programs that always follow the rules and can beat most human players for 30 years! A lot of the hypesters are using LLMs to do things that other types of program or abundant human labour are good at.
Bill, the idea that there is just one kind of intelligence and it is measurable and applies to things other than humans is key to the cult and eugenicist side of this space. Maciej Ceglowski noted that Stephen Hawking was smarter than his cat, but that did not mean he could make his cat do what he wanted. And everyone on a thread like this has known very intelligent people whose lives were disasters.
The linked post from Gary Smith concludes with:
This is wrong. It misunderstands the asymmetry between generating a correct answer and verifying a correct answer. I use LLMs all the time to generate plotnine (Python’s knockoff of ggplot2) and pandas (Python’s knockoff of data frames and Tidyverse manipulation) code all the time. In some sense I know these tools, but I can never remember the exact incantation to put the x axis on a log scale and remove the ticks and labels from the y axis. When I tell the chatbot what I want and it generates pandas and plotting code, I can verify that it’s correct. Doing this used to take me forever as I would have to either investigate the doc or StackOverflow answers or tutorials. Now I put many more figures in things I’m writing.
The asymmetry between verification and generation is key to understanding the difference between polynomial time (P) and non-deterministic polynomial time (NP) algorithms. An NP algorithm can be formulated as guessing with a P algorithm and verifying with a P algorithm. If the LLM is a much better guesser than me, it saves me a huge amount of search. It is also really great at writing sort snippets of code from tight text descriptions. I can say what I want and it can generate what I want faster than I can generate what I want. So it’s a huge win for things like graphing.
Guessing and verifying is also the basis for the LLM “thinking” modes like you see in all the chatbots now (this kind of chain-of-thought came to everyone’s attention with DeepSeek). The drawback is the amount of compute it takes to guess.
One of my main use cases for chatbots is scenario and character generation for roleplaying games. I’m playing a short campaign Fast and Furious knock off now set in Paris in May of ’68. Not only is the chatbot amazing at suggesting locations and plot lines, it’s amazing at inventorying the cars that the Italian, French, German, and British teams would drive in 1968. And then it can generate the tokens. I’ll give you an example in a second comment.
Bob:
That’s interesting, the difference between verifying and generating. I guess that many of the problems we see with chatbots (such as the notorious report with fake references prepared by the US Dept of Health and Human Services) came from people using the chatbot to generate but then not using humans to verify.
Last week I generated and debugged about 1400 lines of Python code with Codex CLI and gpt-5 (reasoning set to medium, later high), expanded from a skeleton of notebooks from a junior programmer. The code itself is somewhat complicated LLM prompting with various APIs and two-level batching and Levenshtein compression to avoid excess prompting costs, and more lines to aid development (dry runs, simulated batch API).
I don’t have much to complain. Bugs were minor, mainly around JSON object formats and robustness from the APIs. The agent mostly debugged the code itself, run trials, generated Excel input files etc. My role was managerial and strategic, although the experience _is_ a bit nerving, even when everything is in a repository so changes cancellable. The feeling comes from finding the right balance between control and efficiency, like how much to give the model autonomy and how that reflects in the end result and its understandability.
I will not have time to read and understand all the details in the code, but I don’t think it’ll easily break in ways that would not be detected. The main risk is losing money (to API costs) if large runs need to be repeated.
(Just an example of that asymmetry. And IMO the whole issue of unrealiability in these discussions is blown to absurd proportions, from the practical point of view. I use LLMs excessively for web search, coding, and just for chatting about interesting things, and don’t see unreliability almost any kind of problem, as long as I use top reasoning models and common sense.)
“If you know the answer, you don’t need to ask an LLM; if you don’t know the answer, you can’t trust one.”
I liked the tic-tac-toe experiment a lot, but the above conclusion… ugh. Somehow it reminds me of this gem,
“Bayesian inference is the only good method of inference if you have a good prior. But if you had a good prior, you wouldn’t need Bayesian inference.”
I wonder whether there is an association between {liking the pragmatism of Bayesian inference} x {understanding that LLM are indeed amazing if you know how to play to their strengths}. Does it follow that fully-skeptical-anti-LLMers are those who firmly believe in the power of “sampling-distribution-theoretical-guarantees-of-optimality means that you have solved your problem, case closed”?
Here’s the image of a sliver-blue ’62 Jaguar XKE convertible it came up with. The initial prompt was:
Here’s what it came up with.
I asked it about cars I might have missed (my friends and I grew up in Detroit and thus car geeks by birth). My friends are playing the British team. They’re WW2 vets going out for “one last score”.
And here’s the reply (removing 8 of the 10 cars listed):
1. Lotus
Lotus Elan S3 (1966–1968)
Lightweight roadster with nimble handling; a favorite of motoring journalists.
Standard colors: Cirrus White, Carnival Red, British Racing Green, Lagoon Blue, Silver Frost
Lotus Cortina Mk1 (1963–1966, with Ford)
High-performance saloon built by Lotus on a Ford platform; dominant in rally and touring car racing.
Standard colors: Ermine White with Sherwood Green stripe
…
10. Rover
Rover P5B Coupe (1967–1973)
V8 coupe favored by government ministers and gangsters alike.
Standard colors: Arden Green, Zircon Blue, Burnt Grey, Silver Birch
These vehicles round out the British fast/performance options available up to 1968, each with distinctive character—ranging from brutal V8 muscle to refined GT elegance to nimble rally engineering. Let me know if you want suggestions sorted by role (e.g., getaway car, street racer, aristocrat’s tourer).
I actually got the main plot line from GPT—there’s a fake Mona Lisa actually hanging in the Louvre—the original was packed off for safe keeping, but never left Gare d’Orsay—it’s just waiting in the abandoned terminal (now an art museum) for a final encounter. But for now, they’re on a mad dash around Paris to find it. They’ve already had their first illegal street race from the Palmarium (now the Orangerie) in the Bois de Boulogne (huge park) to the Arc de Triomphe, three laps around, then back. The German team leader wound up crashing his gull wing Mercedes into the Palmarium in an attempt to beat our party on foot (classic Euro/contintental race style is to start on foot). The after party was at Moulin Rouge (my players thought it may have been fictional!), naturally, but only after being chased by the police after crashing the Palmarium. The party was followed by a midnight gun battle at Père Lachaise to recover the first clue from the crypt of Colonel Fabien (the exact crypt also a GPT suggestion given our heroes are war veterans).
So yes, I’m finding chatbots incredibly useful both for work and for play.
To toss a bone to the haters, it has a really really hard time putting the steering wheel on the right side for UK cars. As a bonus, here’s the film poster it generated, which took about 10 tries to get right.
I’m not at alll surprised it had trouble with left and right.
I’ve encountered that as well, as you’ll see in a recent article that I published recently in 3 Quarks Daily: ChatGPT Makes Images. It’s FUN! (and illuminating). The article shows a number of things, starting out with giving ChatGPT a photo and then asking it to use the photo as a foundation for a new image. That gave us an Indiana Jones adventure that travels to another planet.
Then I asked it to create images of a new superhero, Kama Carnatica (“Carnatic” for South Indian classical music): “I want an image of a new superhero. She’s female and goes into action to advance pleasure, equality, and justice. She wears a magic sari that can change from hot pink flames to cool blue waters as the occasion dictates.”
Next I had it do a four-panel comic for kids: 1) a boy on the beach is confronted by a tsunami, 2) it threatens to drown him unless he guesses its name, 3) the boy does so, and 4) is riding a surf board on the wave.
I decided to go crazy and asked it to imagine some classic Western paintings with South Asian subjects: 1) Botticelli’s Birth of Venus, 2) da Vinci’s Mona Lisa, 3) Whistler’s Mother, and 4) Sargent’s Madame X. It had problems with that one. In Sargent’s painting Madame X is facing to the right (our right, her left). ChatGPT had the parody subject facing left (her right). I asked it to correct the error two times, and two times it repeated it. I then asked it what we knew about the difficulties young children have with left and right.
I finish by having it create mandalas, two based on photographs and one based on a prose statement.
Minor point of order: There are 20,000 people in Japan who would not (I’d guess if they were around to be asked) think jokes about tsunamis to be funny.
Seriously, though.
“I’m not at alll surprised it had trouble with left and right.”
Sure. LLMs aren’t good at left right, rotational symmetry, and the like. We all know that.
But they can solve Math Olympiad problems.
Modern abstract mathematics’ most basic foundations are set theory and group theory. If you can’t do group theory, you can’t do any of mathematics beyond calculus.
But group theory is about dealing with symmetries. Which LLMs can’t do. Yet they get credit for “solving math problems”..
This is seriously ridiculous.
You all might be interested in these followups:
Two more examples:
https://drive.google.com/file/d/1koen5uX2-aRBQWtFd1TEIbFcA99Njxcf/view?usp=share_link
GPT 5.0 struggles to generate and evaluate:
https://drive.google.com/file/d/1Vc2KVULTbw4xUVO50QhpwgkC-E50nXDc/view?usp=sharing
I think this approach misses the point that one can have a productive conversation with an LLM and learn something as a result of that. I don’t really see the value in saying “look the model doesn’t live up to the sales hype in these situations, therefore it’s stupid and useless”. Well of course; the hype is hype; but that doesn’t mean that the product has no useful applications. Bob C makes this point extensively above ofc.
Anon:
I think a good analogy is Google. You can have a productive interaction with a Google search and learn something as a result of that. Setting aside any hype, Google is useful. The problem comes when someone sees something on Google (or Facebook, or Twitter, etc.) and assumes it’s accurate.
Compared to Google, the chatbot can produce a convincing form of conversation and it can also be helpful in all sorts of ways. The problem comes when the convincing conversation leads people to assume the content is correct.
Fair enough, although all ChatGPT conversation screens say “ChatGPT can make mistakes. Please check important information.” at the bottom, whereas there is generally no such warning when reading things found via search engines or on social media. But I see what you mean generally with regards to a chatbot conversation style leading a user to have increased trust by virtue of tone (although also worth noting that tone of response can be set or guided).
But with Google, you know who you are reading. If it’s wiki, it’s probably pretty good. Ditto NIH (well, until recently). Mayo Clinic is tooting their own horn to some extent, but it’s reliable in terms of figuring out what’s current best care.
The text in the link itself has a lot of information, that gets lost in the LLM sausage machine.
Again, if you know who you are reading, you can interpret it. I don’t want a random text generator’s sausage of stuff I don’t know the trustworthiness of. And that’s without the halucinations.
David
I don’t understand your comment about Google (“you know who you are reading”). I find the opposite – with Google I’m never sure how the results were generated, how much has been affected by Google’s revenue flows, and how the algorithms have selected what I am being shown. Much like LLMs. Wikipedia seems different to me. So, can you explain what you mean by this distinction?
Dale,
When you get a selection of links you don’t know how those links were selected. But when you click the link… you know what site you’re going to and who wrote it… at least if it’s not some random blog. But then you know what you’re getting if the site identifies itself as being run by a certain professor at a certain school, or a particular industry personality well known in a certain field, or the govt of Spain or whatever.
Just ask for sources, this interface will do it:
https://www.perplexity.ai/
You have to actually check these sources though. Its still only about 60/40 that the refs actually contain the info, even if the bot gives a direct quote.
But that’s how you should be reading human-generated content with references anyway. I’d guess biomed literature is on the order of 1% false references. But for media, like BBC/NYT/etc, it is near 100% frequency that the academic reference does not contain precisely what is claimed.
Anyway, the solution is to have it give exact quotes with page/line numbers, then go back and verify the ref actually exists and contains the content.
Some things you can check. But, the source of the information helps you evaluate it. Chatbots can’t reason, so have no idea whether what they are telling you is correct. Of course, this may be true of some people you know, too. It is tricky because chatbots are pretty good at giving the appearance that they can reason. I like David-in-Tokyo’s analogy of psychics.
A couple of people I know have been telling me chatbots can help me write code. One said to use Github Copilot. So, I tried it. But, its suggestions were useless/wrong. So the other person told me that Copilot is no good; I should use Claude Code. Haven’t tried that (yet?). I’m willing to believe a chatbot can help you write small self-contained apps. Although, it doesn’t take me long to write such apps. Even there, you have to carefully check the code it produces, since it ignores things you said and doesn’t tell you that it has ignored them (since it can’t reason).
The point of real (not halicinated) links is that the TEXT of the URL itself is /nih/ or /wikipedia/ or whatever, and that’s often a lot of information.