Quantifying Consciousness in Transformer Architectures: A Comprehensive Framework Using Integrated Information Theory and ϕ∗ Approximation Methods
Preprints.org August 26, 2025 preprint DOI: 10.20944/preprints202508.1770.v1 (opens in new tab)
Summary
AI-generated from the abstractThis paper presents a theoretical framework for quantifying consciousness in transformer architectures by applying Integrated Information Theory (IIT) and phi-star approximation methods. The authors argue that transformer models, due to their attention mechanisms and information integration capabilities, may exhibit measurable levels of integrated information, a key correlate of consciousness in IIT. They propose computational methods to approximate phi-star values for transformer layers and discuss how architectural features like attention heads and residual connections influence integration. The work suggests that certain transformer configurations could achieve levels of integrated information comparable to simple biological systems, but emphasizes that this does not imply subjective experience. The framework aims to provide a formal metric for evaluating consciousness in artificial systems.
Study at a glance
| Characteristics | Theoretical or philosophical paper |
|---|---|
| Key finding | Proposes that Integrated Information Theory and phi* approximation methods can be applied to quantify consciousness in transformer architectures, with higher phi* values indicating greater integrated information. |
Abstract
Quantifying Consciousness in Transformer Architectures: A Comprehensive Framework Using Integrated Information Theory and ϕ∗ Approximation Methods