0429 DreamGPT: Validation of ChatGPT as a Tool for Emotion Analysis in Dream Reports
Garrett Baber, Nancy Hamilton, Matthew Gratton
Sleep May 19, 2025 DOI: 10.1093/sleep/zsaf090.0429 (opens in new tab)
Summary
AI-generated from the abstractChatGPT's ratings of positive and negative affect in dream reports showed excellent agreement with participants' own ratings. In 136 dream reports, the model's positive affect ratings matched self-reports with an intra-class correlation coefficient of 0.844 and a mean absolute error of 1.778 on a 0–10 scale. For negative affect, agreement was similarly high, with an ICC of 0.857 and a mean absolute error of 1.681. Bland-Altman plots revealed no systematic bias. The findings suggest that large language models can reliably estimate emotional content in dream narratives, offering a scalable alternative to self-report for sleep and affective research.
Study at a glance
| Characteristics | Observational study Peer reviewed |
|---|---|
| Sample size | 136 |
| Population | Participants who provided one dream report each |
| Key finding | ChatGPT showed excellent agreement with self-reported positive and negative dream affect ratings, with ICCs of 0.844 and 0.857 respectively. |
Abstract
Abstract Introduction Dream reports provide unique insights into emotional processing during sleep. While self-reports are the gold standard for assessing dream affect, natural language processing (NLP) tools like ChatGPT may offer scalable alternatives. This study evaluated ChatGPT’s ability to estimate positive and negative affect from dream reports, comparing its performance against self-reported ratings. Methods A total of 136 participants provided one dream report each. Participants rated their positive and negative dream affect on a 0–10 scale. ChatGPT 3.5 Turbo was accessed via the OpenAI API using Python, where each dream report was analyzed and rated on identical scales. The model’s outputs were programmatically appended to a CSV file for further analysis. Agreement between ChatGPT and self-reports was assessed using intra-class correlation coefficients (ICC3k) for consistency, mean absolute error (MAE) for deviation, and Bland-Altman plots for visual inspection of agreement. Results For positive affect ratings, ChatGPT demonstrated excellent agreement with self-reports (ICC3k = 0.844, 95% CI [0.781, 0.889], p <.001), with an MAE of 1.778. Negative affect ratings similarly showed excellent agreement (ICC3k = 0.857, 95% CI [0.799, 0.898], p <.001), with an MAE of 1.681. Bland-Altman plots for both affective dimensions indicated no systematic bias and acceptable limits of agreement upon visual inspection. Conclusion ChatGPT demonstrated strong agreement with self-reported dream affect ratings, supporting its potential as a scalable tool for analyzing emotional content in dream reports. These findings suggest that large language models can provide valid and reliable estimates of dream affect, which may advance sleep and affective science. Support (if any)