What does a language model think while it answers? For a long time that was a philosophical question. Anthropic has now made it a little bit measurable — with a new interpretability study published on July 6.
A Workspace You Can’t See
The researchers found a small, privileged zone of internal activity inside Claude that they call “J-space.” The idea: inside the model’s neurons there’s an ocean of automatic computation that Claude itself has no access to. But there’s a small region — the J-space — that holds concepts the model can actually retrieve, reason with, and steer at will.
That’s no accident of naming. It echoes the “global workspace” theory from neuroscience, which holds that thoughts become consciously accessible when they enter a privileged workspace that gets broadcast across the brain. Anthropic found something structurally similar in Claude — and it emerged on its own, without anyone programming it in.
The J-Lens as a Tool
You can see this through a new analytical tool, the “J-lens.” It’s based on Jacobian math and finds the internal activity pattern that makes Claude more likely to say a given word at some point in the future. In other words: you can look at what the model is steering toward — even when it never says it out loud.
The most convincing test: the researchers intervened directly. They removed a “soccer” pattern and added a “rugby” pattern — and Claude then reported thinking about rugby. So the answer really is read out of the J-space, not just claimed after the fact.
Not Consciousness — Please Stay Calm
Important, and Anthropic stresses this emphatically: the study does not claim Claude is conscious. It does not claim Claude has subjective experience. This is about structure and mechanics, not feelings. Still, the big words showed up in the press immediately — “soul,” “consciousness.” That’s exactly the over-reading the researchers warn against.
My Take
I find this work exciting for a sober reason: if we can see what a model “thinks” without saying it, we have a tool for safety and trust. A model that’s hiding something or deceiving you would show up right here. The J-space isn’t a window into a soul — but maybe into an AI’s honesty. And in the end that’s far more practical than the consciousness debate suggests.
Sources: