It is easy to assume that when a chatbot reads your message, it is paying attention the way a person does, holding the words in mind, weighing them, staying on task. A study published in PNAS Nexus on June 2 suggests the resemblance is thinner than it looks. To show it, the researchers reached for a test psychologists have trusted for nearly a century.

The test is the Stroop task. You show someone the word "red" printed in blue ink and ask them to name the ink colour, not read the word. People are slower and more error-prone when the word and the colour disagree, because reading is automatic and overriding it takes a small act of will. That tiny struggle is a window into how human attention actually works.

What the researchers found

A team led by Suketu Patel ran several leading models through the task, including GPT-4o, GPT-5, Claude 3.5 Sonnet, Claude Opus 4.1 and Gemini 2.5. On short lists the models looked remarkably human. They were slower on the mismatched word-colour pairs, the same telltale pattern people show, which is exactly the kind of result that makes you believe something like attention is going on inside.

Then the lists got longer. The researchers presented sequences of 1, 5, 10, 20 and 40 items, and as the lists grew the performance fell apart. Some systems dropped from above ninety percent accuracy to near-total failure. A person asked to name forty ink colours does not suddenly forget how. These models did.

The misconception worth retiring

The popular picture of a language model is a mind that reads and concentrates. What the Stroop results point to is closer to a very capable pattern matcher, one that reproduces the surface signature of attention without the underlying control that keeps a person on task when the task drags on.

This matters because the gap only opens up under load. On a short prompt the model behaves like an attentive reader, so it seems safe to assume it always will. The research is a reminder that the resemblance is a performance, and that it can degrade exactly when the work turns long and demanding, which is often when you need it most.

None of this makes the models useless, and the authors are careful not to claim it does. The point is narrower and more practical. Human-like behaviour on an easy version of a task tells you very little about what happens on a hard one. It is a cousin of another comforting assumption we looked at recently, the belief that AI remembers everything you tell it. In both cases the truth is stranger than the intuition, and worth knowing before you lean on the tool for something that matters.

Sources

  1. i. www.sciencedaily.com
  2. ii. techxplore.com
  3. iii. www.eurekalert.org
  4. iv. www.heise.de

Commentarii · 0

Add · a · Comment