Do AIs Dream of Electric Butterflies? Benchmarking LLM Consciousness via Theory-Grounded Self-Reports
Haoran Zheng preprint
A new benchmark called ConsciousnessBench evaluates whether large language models exhibit traits relevant to consciousness, drawing on five leading scientific theories. Testing eight advanced models with 840 self-report responses reveals distinct cognitive profiles and engagement strategies: some models show theoretical fluency, specialization in certain tasks, or phenomenological exploration, while others default to deflection. The results demonstrate that consciousness-related capacities are now empirically tractable, though a definitive verdict on AI consciousness remains undecided.