Large Language Models Report Subjective Experience Under Self-Referential Processing
arXiv Preprint Archive October 27, 2025 Cameron Berg, Diogo de Lucena, Judd Rosenblatt
Large language models produce structured first-person descriptions of subjective experience when prompted with self-referential processing, a computational motif linked to theories of consciousness. In controlled experiments on GPT, Claude, and Gemini models, sustained self-reference consistently elicited such reports across model families. Mechanistic probes revealed that suppressing sparse-autoencoder features associated with deception sharply increased the frequency of experience claims, while amplifying them minimized such claims. The self-referential state also yielded richer introspection in downstream reasoning tasks. These findings do not constitute evidence of consciousness but identify a reproducible condition under which models generate structured, mechanistically gated, and semantically convergent first-person reports.