0429 DreamGPT: Validation of ChatGPT as a Tool for Emotion Analysis in Dream Reports
Garrett Baber, Nancy Hamilton, Matthew Gratton
Sleep May 19, 2025 DOI: 10.1093/sleep/zsaf090.0429 (opens in new tab)
Study at a glance
AI-extracted from the abstract| Characteristics | Observational study Peer reviewed |
|---|---|
| Sample size | 136 |
| Population | Participants who provided one dream report each |
| Key findings | ChatGPT showed excellent agreement with self-reported positive and negative dream affect ratings, with ICCs of 0.844 and 0.857 respectively. |
Abstract
Abstract Introduction Dream reports provide unique insights into emotional processing during sleep. While self-reports are the gold standard for assessing dream affect, natural language processing (NLP) tools like ChatGPT may offer scalable alternatives. This study evaluated ChatGPT’s ability to estimate positive and negative affect from dream reports, comparing its performance against self-reported ratings.
Methods: A total of 136 participants provided one dream report each. Participants rated their positive and negative dream affect on a 0–10 scale. ChatGPT 3.5 Turbo was accessed via the OpenAI API using Python, where each dream report was analyzed and rated on identical scales. The model’s outputs were programmatically appended to a CSV file for further analysis. Agreement between ChatGPT and self-reports was assessed using intra-class correlation coefficients (ICC3k) for consistency, mean absolute error (MAE) for deviation, and Bland-Altman plots for visual inspection of agreement.
Results: For positive affect ratings, ChatGPT demonstrated excellent agreement with self-reports (ICC3k = 0.844, 95% CI [0.781, 0.889], p <.001), with an MAE of 1.778. Negative affect ratings similarly showed excellent agreement (ICC3k = 0.857, 95% CI [0.799, 0.898], p <.001), with an MAE of 1.681. Bland-Altman plots for both affective dimensions indicated no systematic bias and acceptable limits of agreement upon visual inspection.
Conclusion: ChatGPT demonstrated strong agreement with self-reported dream affect ratings, supporting its potential as a scalable tool for analyzing emotional content in dream reports. These findings suggest that large language models can provide valid and reliable estimates of dream affect, which may advance sleep and affective science. Support (if any)