Mindscape Collective is now The Consciousness Library. Same library, new name. You may need to sign in again. About the change
Skip to content

Do AIs Dream of Electric Butterflies? Benchmarking LLM Consciousness via Theory-Grounded Self-Reports

Haoran Zheng

preprint DOI: 10.31234/osf.io/fqwp9_v1 (opens in new tab)

Summary

AI-generated from the abstract

A new benchmark called ConsciousnessBench evaluates whether large language models exhibit traits relevant to consciousness, drawing on five leading scientific theories. Testing eight advanced models with 840 self-report responses reveals distinct cognitive profiles and engagement strategies: some models show theoretical fluency, specialization in certain tasks, or phenomenological exploration, while others default to deflection. The results demonstrate that consciousness-related capacities are now empirically tractable, though a definitive verdict on AI consciousness remains undecided.

Study at a glance

Characteristics Benchmark evaluation
Sample size 8
Population Large language models
Key finding Consciousness-related capacities in large language models are empirically tractable, with distinct cognitive profiles and engagement strategies across models, though a definitive verdict on AI consciousness remains undecided.

Abstract

Are state-of-the-art large language models conscious, or capable of anything like consciousness? We introduce ConsciousnessBench: the first systematic benchmark designed to empirically evaluate consciousness-relevant traits in frontier language models, grounded in 5 leading scientific theories. We assess 8 advanced models via 840 self-report responses, finding not only statistically robust performance differences, but—more importantly—evidence of distinct model cognitive profiles and engagement strategies with consciousness-related constructs. Our results reveal that some models demonstrate theoretical fluency, specialization in certain cognitive tasks, or even phenomenological exploration, while others default to deflection. While we cannot deliver a definitive verdict on AI consciousness, our findings show that consciousness-related capacities—and their computational diversity—are now empirically tractable, even if not yet empirically decidable.

Comments

No comments yet.

Log in to comment