Synthetic Corpus of Emotions for Detection of Depression in Social Networks
摘要
The detection of mental disorders through Natural Language Processing (NLP) has gained significant attention, particularly in the field of depression detection. Despite the availability of various corpora for English, Spanish-language resources remain limited. To address this gap, this study presents a synthetic corpus based on Plutchik’s emotion model, designed to enhance depression detection in Spanish-language social media messages. This work compares three different corpora: the Twitter Sentiment analysis corpus, the LiSSS corpus, and the created synthetic corpus. Each dataset is used to train emotion-detection models, which then generate emotion vectors as input for a multilayer perceptron neural network to classify users as depressed or non-depressed. The results demonstrate the effectiveness of emotion-based features for depression detection and highlight the potential of synthetic corpora. The comparison shows that the synthetic corpus is statistically similar to a corpus generated and labeled by humans. Moreover, the results of the model based on the synthetic corpus perform as well as or better than other corpora labeled by humans.