Ask a chatbot for a source and it may hand you a citation that looks perfect and does not exist. Ask about an obscure person and it may invent a confident biography. The reassuring story people tell themselves is that this is a bug, a rough edge the next release will sand down until the models simply stop making things up. It is a comforting idea. It is also mostly wrong, and understanding why changes how you should use these tools.
Not a glitch, a consequence
A large language model does not look anything up. It predicts the next piece of text, one token at a time, based on patterns learned from its training data. There is no internal database of facts it consults, and no built-in sense of the difference between a true statement and a plausible one. When the model has seen enough examples, the likely next word happens to be correct. When it has not, the model still produces the most likely-sounding continuation, delivered in the same even, confident tone. From the inside, a fact and a fabrication are generated by the same process. That is why the errors are so fluent, and so easy to miss.
OpenAI made a version of this argument in its own research in 2025, pointing out that the way models are trained and scored can actively reward guessing. If a benchmark gives a point for a correct answer and nothing for admitting uncertainty, a model that always ventures a guess will outscore one that sometimes says it does not know. Honesty, under those rules, is a losing strategy. The hallucination is not the model malfunctioning. It is the model doing exactly what it was optimised to do.
Getting better is not the same as getting fixed
None of this means the problem is static. It is genuinely shrinking. Grounding a model in retrieved documents, letting it search as it answers, training it to express uncertainty and checking its work with other models all push the error rate down, sometimes sharply. Newer systems hallucinate less than the ones from two years ago. But reducing a thing is not eliminating it, and the floor is not zero. The same confident fabrication turns up in the flood of AI-written papers carrying invented references, and it is the reason a chatbot makes an unreliable fact-checker.
How to actually live with it
The practical takeaway is not to distrust these tools, which are genuinely useful, but to use them for what they are. Treat a model's output as a confident draft from a fast, widely read assistant who sometimes gets things wrong and never signals when. Verify anything that matters, especially names, numbers, quotes and citations, against a real source. Be most suspicious when the subject is narrow or obscure, which is exactly where the training data thins out and the guessing begins. Waiting for the release that makes verification unnecessary is waiting for something the underlying design does not promise. The models will keep improving. The habit of checking them should not go anywhere.
Sources
- i. openai.com
Commentarii · 0