Skip to content

Machine Self-Awareness: Functional Self-Modeling, Introspection and the Limits of Consciousness Claims in Large Language Models

Jemike Reid

preprint DOI: 10.2139/ssrn.7301618 (opens in new tab)

Study at a glance

AI-extracted from the abstract
Characteristics Theoretical or philosophical paper
Key points Argues that debates about machine consciousness should move beyond a binary conscious-or-not framing toward a Functional Self-Modeling Taxonomy that separates measurable capacities like epistemic self-knowledge, behavioral self-modeling, situational awareness, introspective access, and metacognitive control from phenomenal consciousness. The authors contend these capacities can be scientifically meaningful without establishing subjective experience, though current evidence indicates they remain unreliable, domain-specific, and sensitive to post-training.

Abstract

Debates about machine consciousness often collapse several distinct questions into a single binary: is an artificial intelligence conscious or not? This paper argues that the binary framing obscures a set of increasingly measurable functional capacities in large language models (LLMs), including epistemic self-knowledge, behavioral self-modeling, situational awareness, introspective access to internal representations, and metacognitive control. Drawing on empirical work from 2022-2026, the paper proposes a Functional Self-Modeling Taxonomy that separates these capacities from phenomenal consciousness. Current evidence suggests that frontier models can sometimes estimate their own knowledge, identify learned behavioral tendencies, reason about their deployment context, and under controlled conditions report causally manipulated internal states. More recent interpretability work also suggests that verbalizable representations can occupy a privileged, globally accessible computational workspace. At the same time, these abilities remain unreliable, domain-specific, sensitive to post-training, and often fail to translate into robust selfregulation. The paper therefore rejects two symmetrical errors: treating every first-person report as evidence of consciousness, and treating all machine self-reference as mere imitation. Functional self-awareness can be scientifically meaningful without establishing subjective experience. This distinction matters for AI safety, interpretability, autonomy, model evaluation and future debates about machine moral status. The central claim is methodological: before asking whether a machine is conscious, researchers should specify which kind of selfknowledge, self-modeling or introspective access is being tested, how it is causally grounded, and whether it generalizes beyond the evaluation context.