You have heard the dismissal, probably more than once. A large language model is just autocomplete on steroids. It predicts the next word, one at a time, from patterns in its training data. It does not know anything. The most memorable version of the idea is the phrase "stochastic parrot," coined in a 2021 paper by Emily Bender and colleagues to describe a system that stitches together plausible language without any grasp of meaning.
It is a good line because it starts from something true. These models are trained to predict the next token, and at run time that is mechanically what they do. Anyone selling you a model as a conscious mind is overselling. So the myth here is not the claim that prediction is involved. It is the word "just."
What the dismissal gets right
Language models have no body, no memory of yesterday unless you give them one, and no stake in whether what they say is true. They will state a falsehood with the same fluency as a fact. They can be led into confident nonsense by a badly worded prompt. If your mental model is a knowledgeable friend, these failures look bizarre. If your mental model is a very good next-word predictor, they look exactly like what you would expect. The skeptics are right that fluency is not understanding, and that treating the two as the same causes real mistakes.
Where "just" breaks down
The trouble is that "predict the next word" turns out to be a deceptively deep task. To guess the next word in a murder mystery, or a chess game written in notation, or a paragraph of legal argument, a system has to build some internal representation of what is going on. Researchers who probe inside these models have found structures that look like maps of space, models of a game board, and features that track abstract concepts. The prediction is the training goal. What the model builds in order to get good at it is the interesting part.
As the cognitive scientist and others have put it, calling a modern model "just next-token prediction" is a bit like calling a human brain "just neurons firing." Both statements are technically accurate and both miss the thing worth explaining. Even Bender has spent time correcting how her phrase gets used, noting that the popular version has drifted from the careful argument in the paper.
So which is it
Honestly, the science is not settled, and anyone who tells you it is has picked a side. There is real disagreement among serious researchers about how much genuine understanding these systems have, and the honest answer sits in an uncomfortable middle. A language model is not a mind, and it is not a parrot either. It is a new kind of thing that we do not yet have good words for, which is exactly why we keep reaching for old ones.
The practical takeaway is simpler than the philosophy. Do not trust a model because it sounds sure of itself, and do not dismiss it because you have heard it is only autocomplete. Judge it by what it can and cannot reliably do. That test survives whichever way the deeper argument eventually falls. If you want a related version of this trap, we looked recently at whether a model that thinks out loud is really reasoning, and at why a bigger model is not automatically a smarter one.
Sources
- i. spectrum.ieee.org
- ii. en.wikipedia.org
- iii. medium.com
Commentarii · 0