Automatic Annotation of Dream Report’s Emotional Content with Large Language Models
Lorenzo Bertolini, V. Elce, Adriana Michalak, Hanna-Sophia Widhoezl, Giulio Bernardi, J. Weeds
Workshop on Computational Linguistics and Clinical Psychology January 1, 2024 DOI: 10.18653/v1/2024.clpsych-1.7 (opens in new tab) via Semantic Scholar
Summary
AI-generated from the abstractDream reports are typically analyzed by trained human annotators, a time-consuming process. While earlier natural language processing tools could automate some analysis, they could not reason over full report context, required extensive preprocessing, and were rarely validated against manual scoring. This study used large language models (LLMs) to replicate manual annotation of dream reports, focusing on references to emotions. An off-the-shelf LLM method performed poorly, likely due to linguistic differences between reports from different individuals. A bespoke text classification method achieved high performance and was robust against biases. The approach may enable analysis of large dream datasets and improve reproducibility across studies.
Study at a glance
| Characteristics | Empirical study Peer reviewed |
|---|---|
| Keywords | Computer science Psychology |
| Key finding | A bespoke text classification method using large language models achieved high performance in replicating manual annotation of emotions in dream reports, while an off-the-shelf method performed poorly. |
Abstract
In the field of dream research, the study of dream content typically relies on the analysis of verbal reports provided by dreamers upon awakening from their sleep. This task is classically performed through manual scoring provided by trained annotators, at a great time expense. While a consistent body of work suggests that natural language processing (NLP) tools can support the automatic analysis of dream reports, proposed methods lacked the ability to reason over a report’s full context and required extensive data pre-processing. Furthermore, in most cases, these methods were not validated against standard manual scoring approaches. In this work, we address these limitations by adopting large language models (LLMs) to study and replicate the manual annotation of dream reports, using a mixture of off-the-shelf and bespoke approaches, with a focus on references to reports’ emotions. Our results show that the off-the-shelf method achieves a low performance probably in light of inherent linguistic differences between reports collected in different (groups of) individuals. On the other hand, the proposed bespoke text classification method achieves a high performance, which is robust against potential biases. Overall, these observations indicate that our approach could find application in the analysis of large dream datasets and may favour reproducibility and comparability of results across studies.