Mindscape Collective is now The Consciousness Library. Same library, new name. You may need to sign in again. About the change
Skip to content

Probing the Representational Geometry of Color Qualia: Dissociating Pure Perception from Task Demands in Brains and AI Models

Jing Xu

arXiv Preprint Archive October 26, 2025 via arXiv

Summary

AI-generated from the abstract

The representational geometry of color qualia differs between AI vision models and the human brain. Using fMRI with a no-report paradigm, the study compared neural activity during pure perception versus task-modulated perception against diverse vision models. Most models aligned better with pure perception, indicating current feedforward architectures do not capture cognitive processes during task execution. Training paradigm and architecture interact critically: Contrastive Language-Image Pre-training (CLIP) improved brain-alignment for a vision transformer but worsened it for a ConvNet. The work provides a new benchmark for color qualia, revealing fundamental divergence in inductive biases between artificial and biological vision systems.

Study at a glance

Characteristics Observational study with fMRI and Representational Similarity Analysis Peer reviewed
Population Human participants in fMRI study
Keywords Cs.ne
Key finding Nearly all AI vision models align better with neural representations of pure perception than with task-modulated perception, and multi-modal training (CLIP) improves brain-alignment for vision transformers but harms it for ConvNets.

Abstract

Probing the computational underpinnings of subjective experience, or qualia, remains a central challenge in cognitive neuroscience. This project tackles this question by performing a rigorous comparison of the representational geometry of color qualia between state-of-the-art AI models and the human brain. Using a unique fMRI dataset with a "no-report" paradigm, we use Representational Similarity Analysis (RSA) to compare diverse vision models against neural activity under two conditions: pure perception ("no-report") and task-modulated perception ("report"). Our analysis yields three principal findings. First, nearly all models align better with neural representations of pure perception, suggesting that the cognitive processes involved in task execution are not captured by current feedforward architectures. Second, our analysis reveals a critical interaction between training paradigm and architecture, challenging the simple assumption that Contrastive Language-Image Pre-training(CLIP) training universally improves neural plausibility. In our direct comparison, this multi-modal training method enhanced brain-alignment for a vision transformer(ViT), yet had the opposite effect on a ConvNet. Our work contributes a new benchmark task for color qualia to the field, packaged in a Brain-Score compatible format. This benchmark reveals a fundamental divergence in the inductive biases of artificial and biological vision systems, offering clear guidance for developing more neurally plausible models.

Comments

No comments yet.

Log in to comment