Comparative study of LSA vs Word2vec embeddings in small corpora: a case study in dreams database
Edgar Altszyler, Mariano Sigman, Sidarta Ribeiro, Diego Fernández Slezak
arXiv Preprint Archive October 5, 2016 via arXiv
Summary
AI-generated from the abstractWord embeddings are typically studied in large text datasets, but few studies examine small corpora like single-person text production. This paper compares Skip-gram and LSA for extracting semantic patterns from small dream report series. LSA outperformed Skip-gram on small training corpora in two semantic tests. As a case study, LSA captured relevant word associations in dream reports even with few dreams or low-frequency words. The authors propose LSA can explore word associations in dream reports, offering new insights for psychology research.
Study at a glance
Abstract
Word embeddings have been extensively studied in large text datasets. However, only a few studies analyze semantic representations of small corpora, particularly relevant in single-person text production studies. In the present paper, we compare Skip-gram and LSA capabilities in this scenario, and we test both techniques to extract relevant semantic patterns in single-series dreams reports. LSA showed better performance than Skip-gram in small size training corpus in two semantic tests. As a study case, we show that LSA can capture relevant words associations in dream reports series, even in cases of small number of dreams or low-frequency words. We propose that LSA can be used to explore words associations in dreams reports, which could bring new insight into this classic research area of psychology