Ask a chatbot for a citation and it may invent one, complete with plausible authors and a journal that does not exist. The common reaction is to call the model a liar, or to imagine some glitch where the machine briefly loses its grip on reality. Both readings assume there is a truth the model knows and is choosing to withhold or scramble. That is the misconception worth taking apart, because the evidence points somewhere far more mundane.
A language model has no beliefs to betray. It generates text by predicting, one piece at a time, what is statistically likely to come next given everything before it. When the honest answer would be "I am not sure," the model has usually not been given any reason to say so. It produces the most probable-looking continuation instead, and a confident fabrication often looks more probable than a hedge. There is no intent to deceive because there is no intent at all.
Why guessing gets rewarded
The more useful question is why models are built this way, and here recent research is unusually direct. In a 2025 paper titled Why Language Models Hallucinate, researchers at OpenAI argued that standard training and testing procedures actively reward guessing. Think of a student sitting a multiple-choice exam. If a wrong answer scores the same as a blank, but a lucky guess might score a point, the rational move is always to guess. Models face the same incentive at scale.
The team looked at the popular benchmarks used to rank AI systems and found that most of them grade on accuracy alone. A model that admits uncertainty is scored exactly like a model that is confidently wrong, so the one that guesses wins the leaderboard. A separate analysis published in Nature reached a similar conclusion: evaluating models purely for accuracy incentivizes the very behaviour everyone complains about.
What that means for the fix
If the cause is an incentive rather than a defect in reasoning, the remedy follows. The researchers propose changing how models are scored, penalising a confident error more heavily than an honest "I don't know," and giving partial credit for well-calibrated uncertainty. That is less about rebuilding the model and more about rebuilding the exam it is trained to pass. It will not drive fabrication to zero, but it changes what the system is rewarded for saying.
None of this makes the practical risk go away. A made-up case citation is just as useless whether it came from malice or from math, which is why the sensible habit is still to verify anything a model tells you that matters. What the framing changes is where you look for a solution. Treating hallucination as lying invites the wrong fixes, like scolding the model or hunting for hidden motives. Seeing it as a predictable product of training points at the thing you can actually adjust. The same clear-eyed caution applies to related claims, from tools that promise to detect AI writing to confident numbers about how much faster AI makes people work. The machine is not out to fool you. It is doing exactly what it was rewarded to do.
Sources
- i. openai.com
- ii. arxiv.org
- iii. www.nature.com
Commentarii · 0