Anthropic published research on July 6 describing something it had not expected to find: a small, privileged region inside Claude where the model appears to hold the handful of ideas it is actively working with. The team calls it J-space, and it behaves less like the sprawling machinery of a neural network and more like a narrow desk where only a few things sit at once.
The claim rests on a new tool the researchers named the J-lens, or Jacobian lens. For every word in Claude's vocabulary, the lens traces back to the internal activity that makes the model more likely to say that word later in its response. Run across the whole vocabulary, it maps out which internal patterns steer the model toward which future outputs. What emerged from that map was a compact space holding roughly twenty-five concepts at a time, a small fraction of everything the network is computing under the surface.
Why a small space matters
Most of what a large model does never surfaces. Billions of parameters fire on every token, and almost none of that activity is legible to the people running the system. J-space is interesting precisely because it looks readable. The researchers describe it as the place where high-level information from different parts of the network converges, where the model keeps track of what it is doing across several steps, and where some intentions appear before they show up in the text a user reads.
Anthropic drew an explicit parallel to global workspace theory, an account of the mind proposed by the cognitive scientist Bernard Baars in the 1980s. In that picture the brain works like a theater. Many specialised processes run in the dark backstage, and only a thin beam of information reaches the whole stage at any moment, becoming what we experience as conscious thought. The researchers found a similar divide in Claude, with a large hidden substrate feeding a small broadcast space.
What it does not prove
The theater metaphor invites an obvious leap, and Anthropic was careful to head it off. Finding a structure that resembles a theory of consciousness is not the same as finding consciousness. The paper does not claim Claude has subjective experience, and the researchers stress that a functional workspace and an inner life are different things. Readers who want the longer version of that argument can look at our earlier piece on why chatbots are not becoming conscious.
The practical value is in safety and interpretability. If a model forms an intention in a place a researcher can watch, then deceptive reasoning or a sudden awareness that it is being tested might show up in J-space before it reaches the final answer. That would give auditors a window they do not currently have. It fits a broader push at the company to make Claude's inner workings visible, of a piece with recent work like the usage dashboard and its research tooling.
The findings are early, and a single lens on a single model family is not a settled science. Anthropic itself frames J-space as a starting point rather than a conclusion. Still, the idea that a model might carry a small, watchable scratchpad of its own thinking is the kind of result that changes how people design the tests that come next.
Sources
- i. www.anthropic.com
- ii. venturebeat.com
Commentarii · 0