Can “consciousness” be observed from large language model (LLM) internal states? Dissecting LLM representations obtained from Theory of Mind test with Integrated Information Theory and Span Representation analysis
Natural Language Processing Journal June 27, 2025 Jingkai Li
Integrated Information Theory (IIT) offers a quantitative framework for explaining consciousness, proposing that conscious systems consist of elements integrated through causal properties. This study applies IIT 3.0 and 4.0 to sequences of Large Language Model (LLM) representations, analyzing data from existing Theory of Mind (ToM) test results. It investigates whether differences in ToM test performances, as presented in LLM representations, can be revealed by IIT estimates such as Φmax, Φ, Conceptual Information, and Φ-structure. These metrics are compared with Span Representations independent of consciousness estimates to differentiate potential consciousness phenomena from inherent separations in representational space. Results indicate that sequences of contemporary Transformer-based LLM representations lack statistically significant indicators of observed consciousness phenomena but show intriguing patterns under spatio-permutational analyses.