Sentient AI in robots and agents: prolegomena for an evidence-based research program
Frontiers in Psychology August 19, 2026 DOI: 10.3389/fpsyg.2026.1903644 (opens in new tab)
Study at a glance
AI-extracted from the abstract| Characteristics | Theoretical or philosophical paper Preregistered Peer reviewed |
|---|---|
| Keywords | Ai ethics Ai sentience Artificial consciousness Cognitive architectures Human-robot interaction Mechanistic interpretability |
| Key points | Proposes a prolegomenal framework for sentient AI research that integrates conceptual disambiguation, multi-theory indicator profiles, causal-mechanistic testing, and robotics-specific evidence and governance. Argues for domain-specific ordinal evidence levels rather than binary verdicts or an aggregate sentience score, and outlines welfare- and valence-relevant tests, a preregistered rating procedure, and a protocol for an embodied care robot using sensorimotor lesions, self-location manipulations, memory ablations, and anti-anthropomorphism controls. |
Abstract
The possibility of sentient artificial intelligence has moved from speculative philosophy to a practical interdisciplinary problem for AI, robotics, and human-robot interaction. Large language models, multimodal agents, and embodied robots can now produce first-person reports, maintain dialogue, use tools, act through sensors and effectors, and participate in socially meaningful contexts. These capacities invite two symmetrical errors: anthropomorphic over-attribution and premature dismissal. This Perspective proposes a prolegomenal framework for future research on sentient AI. Its distinctive contribution lies in operationally integrating four elements that have largely been developed in separate literatures: conceptual disambiguation, multi-theory indicator profiles, causal-mechanistic testing, and robotics-specific evidence and governance. The paper distinguishes sentience, consciousness, self-modeling, metacognition, agency, moral patienthood, and AI welfare; separates evidence about an AI system from evidence about human attribution; and proposes domain-specific ordinal evidence levels rather than binary verdicts or an aggregate sentience score. It further specifies welfare- and valence-relevant tests, a preregistered rating procedure, and a concrete protocol for an embodied care robot using sensorimotor lesions, self-location manipulations, memory ablations, and anti-anthropomorphism controls. The aim is not to offer a definitive test for machine sentience, but to show how research could become more scientifically tractable, psychologically informed, robotics-relevant, and ethically responsible.