Watch a modern model solve a hard problem and it will show its working, laying out a chain of thought step by step before it lands on an answer. It reads like a person reasoning aloud, and that resemblance has hardened into a common belief: that these systems now think the way we do. The evidence for that belief is shakier than the transcripts make it look.

The useful part is real. Prompting a model to work step by step does improve its accuracy on maths and logic, which is why the technique is everywhere. The question is whether the written steps are the actual cause of the right answer, or a plausible story told after the fact. A study from researchers at Northeastern and UC Berkeley found that somewhere between 30 and 60 percent of the reasoning steps models produce had little measurable effect on the final answer. The model reached its conclusion by another route and narrated a tidy one.

Right answer, wrong reasons

Quanta Magazine laid out the puzzle plainly in July: models often get the right answer for reasons that have nothing to do with the explanation they give. Underneath, the machinery is still next token prediction. The chain of thought is text the model generates because that kind of text tends to precede correct answers in its training, not a window into a deliberating mind.

This is not a reason to dismiss the systems, and it is not a claim that they are useless. It is a caution about reading too much into the transcript. There is a genuine safety argument, made in recent work on chain of thought monitorability, that being able to watch a model's steps is valuable precisely because we might catch bad behaviour in the act. But the same researchers warn that this visibility is fragile, and that a legible chain of thought is not proof the model is doing what it says.

Why the distinction matters

Confusing a convincing explanation with genuine understanding is how hype gets made. It is the same leap that lets a model's success on competition maths get read as out-thinking mathematicians, and it feeds the sense that superintelligence is just around the corner. A model that writes a persuasive rationale is doing something impressive. It is not, on the current evidence, thinking in the way the word usually means. Holding both of those ideas at once is the honest position.

Sources

  1. i. www.quantamagazine.org
  2. ii. arxiv.org
  3. iii. www.seangoedecke.com

Commentarii · 0

Add · a · Comment