A large corpus of lucid and non-lucid dream reports
arXiv (Cornell University) March 27, 2026 DOI: 10.48550/arxiv.2603.26992 (opens in new tab) via OpenAlex
Summary
AI-generated from the abstractLucid dreams—dreams in which the dreamer is aware they are dreaming—are difficult to study because they are rare and hard to induce, leaving their characteristics unclear. A large corpus of 55,000 dream reports from 5,000 contributors was assembled by scraping ten years of publicly available anonymous dream journals from an online forum. Users optionally labeled their dreams as lucid, non-lucid, or nightmare, providing 10,000 lucid, 25,000 non-lucid, and 2,000 nightmare labels. Analysis confirmed that language patterns in lucid-labeled reports match known features of lucid dreams. This corpus supports broad dream research, and the labeled subset enables new discoveries about lucid dreaming.
Study at a glance
| Characteristics | Observational study Peer reviewed |
|---|---|
| Sample size | 5,000 |
| Population | Contributors to an online forum who shared anonymous dream journals |
| Topics | Dreaming Lucid dreaming |
| Keywords | Clarity Categorization Construct python library |
| Key finding | Language patterns in lucid-labeled dream reports are consistent with known characteristics of lucid dreams. |
Abstract
All varieties of dreaming remain a mystery. Lucid dreams in particular, or those characterized by awareness of the dream, are notoriously difficult to study. Their scarce prevalence and resistance to deliberate induction make it difficult to obtain a sizeable corpus of lucid dream reports. The consequent lack of clarity around lucid dream phenomenology has left the many purported applications of lucidity under-realized. Here, a large corpus of 55k dream reports from 5k contributors is curated, described, and validated for future research. Ten years of publicly available dream reports were scraped from an online forum where users share anonymous dream journals. Importantly, users optionally categorize their dream as lucid, non-lucid, or a nightmare, offering a user-provided labeling system that includes 10k lucid and 25k non-lucid, and 2k nightmare labels. After characterizing the corpus with descriptive statistics and visualizations, construct validation shows that language patterns in lucid-labeled reports are consistent with known characteristics of lucid dreams. While the entire corpus has broad value for dream science, the labeled subset is particularly powerful for new discoveries in lucid dream studies.