Measurement Validation Reveals Five Dissociable Operational Constructs Underlying Self-Model and Agency in Artificial Neural Architectures
Zenodo (CERN European Organization for Nuclear Research) May 25, 2026 DOI: 10.5281/zenodo.20372254 (opens in new tab) via OpenAlex
Summary
AI-generated from the abstractOperational proxies used to measure consciousness-like properties in artificial neural networks often go untested against independent behavioral measures. When tested, a thalamus-inspired agency proxy was negatively correlated with independent agency tests (r=-0.860), and distributed self tests showed poor convergence. Diagnostic experiments revealed the agency proxy primarily measured gating rather than action-outcome control, and self-model scores conflated boundary maintenance, identity persistence, and ownership. After refinement, five operational constructs showed strong convergence with behavioral tests, including action agency (r=0.874) and boundary self (r=0.940). Mechanistic manipulations separated their computational substrates. The pattern argues that proxy validation should precede mechanistic interpretation, and diagnostic failures can refine the construct set.
Study at a glance
| Characteristics | Experimental validation and diagnostic study Peer reviewed |
|---|---|
| Population | Small custom artificial neural network simulations |
| Keywords | Proxy statistics Agency philosophy Construct python library Action physics Matching statistics |
| Key finding | Operational proxies for consciousness-like properties in artificial neural networks often fail independent behavioral validation, but can be refined into constructs with strong proxy-test convergence. |
Abstract
Operational proxies are common in artificial consciousness research, but their validity is rarely tested against independent behavioral measures. We evaluate such proxies in artificial neural architectures by comparing internal-state measures with behavior-level probes. Initial validation exposed substantial mismatches: a thalamus-inspired agency proxy was negatively correlated with independent agency tests (r=-0.860), and generic distributed self tests had poor convergence. Diagnostic experiments showed that the agency proxy primarily measured gating rather than action-outcome control, and that broad self-model scores conflated boundary maintenance, identity persistence, and ownership. After refinement, five operational constructs showed strong proxy-test convergence, including action agency (r=0.874) and boundary self (r=0.940). Mechanistic manipulations then separated their computational substrates: workspace mechanisms supported boundary and identity-temporal measures, action-outcome loops supported agency and action ownership, and meta-monitoring supported distributed body-schema measures. A subsequent minimal attention-based diagnostic showed that action-loop effects generalize to an attention-based substrate, while boundary-self validation remains measurement-limited. All architectures studied here are small custom simulations used as controlled measurement testbeds, not production-scale language models. The pattern argues that proxy validation should precede mechanistic interpretation, and that diagnostic failures of theoretically motivated proxies can themselves refine the construct set.