Dreams are more “predictable” than you think
Lorenzo Bertolini, Sergio Consoli, Julie Weeds
Frontiers in Sleep July 23, 2025 DOI: 10.3389/frsle.2025.1625185 (opens in new tab) via OpenAlex
Summary
AI-generated from the abstractDream reports are easier for large language models to predict than Wikipedia articles, as measured by lower perplexity scores, indicating that dream content is less 'surprising' to these models than general web text. The models also detected differences in dream reports based on gender, visual impairment, and clinical status, mirroring patterns found in prior research. This suggests that machine learning tools can effectively model dream narratives and may capture subtle group-level variations.
Study at a glance
| Characteristics | Observational study Peer reviewed |
|---|---|
| Population | Dream reports from DreamBank and Wikipedia articles |
| Keywords | Psychology |
| Citations | 2 |
| Key finding | Dream reports had significantly lower perplexity scores than Wikipedia articles, and LLMs detected group differences in dream reports based on gender, visual impairment, and clinical status. |
Abstract
Introduction: A growing body of work has used machine learning and AI tools to analyse dream reports, and compare them to other textual content. Since these tools are usually trained on text from the web, researchers have speculated they might not be suited to model dreams reports, often labeled as "unusual" and "bizarre" content. Methods: We used a set of large language models (LLMs) to encode dream reports from DreamBank and Wikipedia. To estimate the ability of LLMs to model and predict textual reports we adopted perplexity, a measure based on entropy, formally, the exponentiated log-likelihood of a sequence. Intuitively, perplexity indicates how "surprising" a sequence of words is to a model. Results: In most models, perplexity scores for dream reports were significantly lower than those for Wikipedia articles. Moreover, we found that perplexity scores were significantly different in reports produced by male vs female participants, and between blind and normally sighted individuals. In one case, we found this difference to be significant between clinical and healthy subjects. Discussion: Dream reports were found to be generally easier to model and predict than Wikipedia articles. LLMs were also found to implicitly encode group differences previously observed in the literature based on gender, visual impairment, and clinical population.