Not Emergent by Accident: The Implicit Global Workspace as a Consequence of Transformer Optimization
Study at a glance
AI-extracted from the abstract| Characteristics | Theoretical or philosophical paper |
|---|---|
| Key points | Argues that a global-workspace-like structure emerges in Transformers as a predictable consequence of next-token training, formalized as a workspace-pressure theorem. Proposes that reusable latent variables are pressured to occupy output-aligned, low-interference directions, creating a low-effective-dimensional subspace for upstream and downstream computations. |
Abstract
Recent interpretability work argues that large language models contain a privileged set of verbalizable internal representations that function like a global workspace. This work develops a principled explanation for why such a structure should be expected to emerge in standard autoregressive Transformers. The central claim is not that Transformers literally implement the biological global workspace architecture, nor that workspace-like representations imply consciousness but how under next-token cross-entropy training, reusable latent variables that affect many future continuations receive stronger optimization pressure to occupy output-aligned, low-interference directions in the shared residual stream. Superposition limits the number of clean independent features, attention creates low-effective-rank communication pressure, and the unembedding matrix supplies an output-aligned gradient sink. Together these forces bias the model toward an implicit workspace which can be categorized as a low-effective-dimensional subspace that upstream computations write into and downstream computations read from. We formalize this as a workspace-pressure theorem using the pullback of the output Fisher geometry through the layer-to-logit Jacobian. The resulting view treats the global workspace as a predictable consequence of the Transformer topology and objective.