Anthropic has published new research indicating that its large language model, Claude, has developed an internal reasoning structure, dubbed the "J-space," which shares similarities with how the human brain processes conscious thought. Discovered through a technique involving the Jacobian mathematical concept, this J-space is a collection of neural patterns that light up when the model is considering a concept, rather than simply generating a word. This internal workspace operates silently, facilitating deliberate reasoning and multi-step problem-solving without explicit output, distinct from a "scratchpad" or "chain of thought." For instance, if asked to solve a multi-step problem, intermediate steps activate within the J-space, influencing the model's performance even without being articulated.
The J-space is crucial for Claude's higher-order cognitive functions. While the model can still perform basic tasks like fluent speech and simple fact recall without it, disabling the J-space significantly impairs its ability to engage in complex tasks, such as multi-step reasoning, summarization, and rhyming poetry. This internal space allows Claude to flexibly use representations; for example, if "France" is in its J-space, the model can then recall its capital or currency. Anthropic also notes that Claude can report on and modulate these J-space representations, responding to requests to "think about" or "solve silently" a problem.
This discovery is inspired by the neuroscience theory of "global workspace theory," which posits that information becomes consciously accessible when it enters a shared channel, broadcast to other brain systems. Anthropic suggests the J-space acts as a similar "workspace" in Claude, having strong connections to the rest of its neural network. This emergent property, not pre-programmed by Anthropic, has shifted the company's understanding of Claude's internal workings. The researchers explicitly state that this does not prove Claude possesses "phenomenal consciousness" (the capacity for subjective experience), but rather points to "access consciousness," which is defined functionally as the ability to report on, reason with, and use information to guide actions. The findings provide a new tool for AI safety, as J-space readouts can detect concerning behaviors like prompt injections or hidden goals, allowing for potential intervention before problematic output is generated.