A recurring claim in casual conversation about AI is that "the chatbot got it right when I asked, so it's basically a search engine that talks." A recent industry study cited by Upwork and others put the error rate of consumer chatbots at roughly one in four answers. That number is high enough to cause real damage when the model is being used as a substitute for a doctor, a lawyer or a fact-checker. It is also a number that most users do not feel, because of a quirk in how large language models present information.

Tone is the trap

The quirk is tone. A language model has no concept of certainty in the way a human does, but it has been trained to produce text in a register that reads as confident. When the model is right, it says so calmly. When the model is wrong, it says so just as calmly. The accuracy of an answer and the confidence with which it is delivered have come unbundled in a way that is genuinely new for written information. A 1995 encyclopaedia entry could be wrong, but the writer at least believed it. A chatbot has no belief either way.

This asymmetry is the heart of the misconception. Researchers call the wrong-but-confident outputs "hallucinations," and the term has become a catch-all for everything from a made-up citation to a misremembered date to a plausible but invented historical event. Across studies summarised by CTO Magazine and the Bipartisan Policy Center, the pattern is consistent. Hallucinations cluster in areas where the training data is thin, where the question is unusual, or where the model is asked to produce specific factual claims (names, dates, numerical figures, quotations). They are rare in genuinely common knowledge.

Two different questions, one delivery style

What the misconception flattens is the distinction between the two kinds of question. A chatbot asked to summarise the plot of Hamlet will reliably get the broad strokes right because the training data contains a thousand summaries of Hamlet. A chatbot asked to summarise a 2026 legal ruling in a specific jurisdiction may invent citations entirely, because the training data did not include the ruling and the model has no mechanism to refuse the question with the same authority it used to answer it.

There is a usable fix. Learn where the cliff edges are. Treat any model output that contains a specific name, date, statute, citation or numerical figure as a claim that needs independent checking before it leaves your hands. Treat broad explanations of well-established subjects as roughly as trustworthy as a competent Wikipedia article, which is to say useful, mostly accurate, and worth verifying when the stakes are high. The myth that chatbots are either oracles or stochastic parrots papers over what they actually are: useful, fluent, occasionally wrong, and unable to tell you which mode they are in right now.

Sources

  1. i. www.upwork.com
  2. ii. ctomagazine.com
  3. iii. bipartisanpolicy.org
  4. iv. www.blueprism.com
  5. v. www.tomsguide.com

Commentarii · 0

Add · a · Comment