As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words. Most of this processing is invisible to you. But some of what takes place in your brain you do have access to—an image that pops into your head, or a deliberate plan you make about where to go shopping. Neuroscientists and philosophers sometimes refer to the latter type of brain activity as “consciously accessible,” to distinguish it from all the other processing that goes on unconsciously. This activity has special properties: we can describe it, control it, and use it for deliberate reasoning, in contrast to all the automatic processing that goes on without our awareness.
In a new paper, we present evidence that a similar distinction has emerged in modern language models like Claude. We find that Claude has developed a small collection of internal neural patterns that, compared to all its other internal processing, play a special role.
We call the collection of these patterns the J-space—named after the technique we used to find them, involving a mathematical concept called the Jacobian. Each J-space pattern is linked to a particular word. But when one of these patterns lights up, it doesn’t mean the model is saying that word—just that the word is on its mind. If you’ve heard of language models having a “scratchpad” or “chain of thought”—text they write to themselves while reasoning—the J-space is something different. It operates silently, in the model’s internal neural activations, allowing the model to think about a concept without writing it down. Notably, the J-space wasn’t designed or programmed by us, but instead emerged on its own during Claude’s training process.
Maybe I am far from the AI research space, but conflating the statistical processes going on in a LLM with terms like ‘mind’, ‘chain-of-thought’, ‘internal neural activations’, and even worse “… is saying the word, just that the word is on it’s mind” makes me feel that this is a buzzword hype-piece rather than an actual paper
There is a very strong effort to anthropomorphise the model in this abstract, which should be a very big no-no for any selfrespecting unbiased researcher.
Turing’s insight is that since consciousness is internal, you should focus on what is observable, not what it “is”
Yep. I’m not an AI expert and it’s been a long time since I studied linear algebra. But something like 90% of what’s reported here strikes me as a very motivated interpretation of basic word association. Which is just what LLMs do, AFAIK.
The Jacobian might be a useful tool for understanding what an LLM is about to do next. But “when we mess with the computer’s internal state, its output changes” is the biggest duh-headslap. That’s not evidence for anything in particular.
I just want to say… what the fuckkkkkk.
Shit has gotten hot so quickly in the AI space in the last 6 months.



