Verbalizable Representations Form a Global Workspace in Language Models
Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous functional distinction has emerged in large language models. Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing. These representations, which we collectively call the J-space, exhibit the functional properties characteristic of a global workspace: their contents can be reported, deliberately summoned and held, used to carry the intermediate steps of silent reasoning, and passed as arguments to arbitrary downstream computations, while automatic processing such as text parsing and routine inference proceeds without them. The J-space also has structural signatures that global workspace theory associates with conscious access: it carries coherent content only in an intermediate band of layers, holds on the order of tens of concepts at a time, and is broadcast by the model’s weights more widely than other representations. These properties make it a practical window into a model’s unspoken thinking.
Introduction. If the mind is an ocean, we spend our lives floating at the surface. Beneath us, an enormous amount of processing takes place without our knowledge: our visual systems parsing the contours of a face, our motor circuits maintaining our posture. At any given moment, only a small fraction of this neural activity is accessible to us. Yet it is this privileged sliver of activity that we rely on to reason deliberately: to plan what ingredients to buy for a recipe, or to puzzle out why an engine won’t start. Such thoughts can be articulated out loud, deliberately held in mind, and brought to bear on whatever task the moment demands. This distinction, between our accessible thoughts and our unconscious processing, is perhaps the most striking feature of human cognition. In this paper, we present evidence that an analogous functional distinction has emerged in modern AI models.
Discussion / Conclusion. The auditing case studies of §5 applied the J-lens to evaluation scenarios and found that the model’s strategic and situational assessments appear in the J-space even when not visible (or prior to being visible) in the model’s output. The model-organism studies found, further, that a misaligned disposition implanted by training leaves a standing signature in the J-space at the start of the model’s responses, on prompts where the model’s surface behavior is unremarkable. These results suggest that the J-lens could be highly useful for alignment monitoring. If a model’s strategic deliberation routes through the J-space, then inspecting the J-space at the positions where decisions are made will reveal that deliberation, and one can monitor for it. As a practical matter, the lens readout is cheap to compute (a single matrix multiplication per layer, with the matrix computed once per model), requires no auxiliary training, and produces output a human can read directly. It can therefore be easily applied at scale to flag transcripts for review.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
When should tasks involve human-AI partnership versus full automation?- Why can't users and AI articulate shared goals together?
- What tasks do users actually want AI to handle versus what can it automate?
- How should designers make invisible AI state legible to users?
- What makes evaluation easier than envisioning for users?
- What stops AI from helping users articulate preferences they cannot express?
- What separates performative behavioral change from actual capability development in AI?
- Can users articulate what they want before AI helps them discover it?
- How do users fail to articulate what they actually want?
- Can prompt engineering overcome the gulf between user intent and AI interpretation?
- Can users articulate their intent before exploring what an AI system finds?
- Why do AI models treat user intent as binary rather than evolving?
- Why are less experienced thinkers more vulnerable to false AI credibility?
- Does AI knowledge precede actual expertise in hyperreal production?
- How does AI substitute polished style for actual expert judgment?
- Why do intellectual products gain false authority from AI-generated form?
- How does AI presentation authority substitute for actual expert judgment?
- How does AI-assisted learning create the Knowledge Custodian paradox in practice?
- What role shifts occur when experts become custodians of AI knowledge?