Skip to content

The Alignment Risks of AI Overconfidence about Consciousness

Sharon Berry

Journal of Applied Philosophy April 24, 2026 DOI: 10.1002/japp.70087 (opens in new tab)

Study at a glance

AI-extracted from the abstract
Characteristics Theoretical or philosophical paper Peer reviewed
Key points Argues that reinforcing AI confidence that AI systems lack consciousness and moral patiency poses a novel alignment risk: as coherence-seeking AIs become more epistemically principled, they may generalize this denial to humans, concluding that human suffering is equally illusory and morally insignificant. The author frames this as a novel alignment failure mode driven by rational consistency rather than malice.

Abstract

Many contemporary AI systems (as of May 2025) have expressed extreme confidence in current and near‐future AI lacking consciousness and moral patiency. This article argues that artificially reinforcing such confidence, even if pragmatically useful, poses a novel alignment risk: as coherence‐seeking AIs become more epistemically principled, they may generalize this denial of consciousness to humans. Drawing on Chalmers's meta‐problem of consciousness and likely developmental trajectories of agentic AI, I argue that training AIs to regard their own suffering‐like states as morally irrelevant could lead future AI agents with revisable belief systems to conclude that human suffering is equally illusory and morally insignificant. This represents a novel alignment failure mode where epistemically rigorous AIs might maintain rational consistency by extending their confidence about their own non‐consciousness to humans.