Enhanced Depression Detection on Social Media Using Advanced Machine Learning and Linguistic Analysis Techniques
摘要
Mental health condition are represented by constant feelings of sadness, hopelessness, and lack of interest in activities once enjoyed by individuals suffering from depression. In this digital era, people turn to social media to express emotions and narrate their daily experiences. Consequently, social media platforms have become valuable tools for appropriate diagnosis of physical and mental state. It has aimed to determine at what extent machine learning can understand the symptoms of depression among social media users by using 20,000 labeled English tweets dataset. TF-IDF features are included along with sentiment polarity, features from the LIWC dictionary and several hybrid features. This study focuses on Multinomial Naive Bayes, logistic regression, and Support Vector Machine models of analysis. In our study, we noticed a substantial increase in the accurate depression detection of using the ML methodologies. Our models provide improvement in the accuracy over the previous research work. The Multinomial Naive Bayes model with hybrid features such as TF-IDF and Sentiment Polarity attained an accuracy of 88.10%, which has been pointed out as a significant improvement in the efficiency of depression detection algorithms. We also attempted this on the subset of the random 4000 tweets and LIWC features where these features acted as binary labels. This simplified approach was still able to obtain the highest value of F1-Score and Accuracy, suggesting that the models bear a high level of depression indicators detection. Adding to the current literature on detecting depression on social media, this experiment indicates that directly labeling LIWC categories might be a useful approach to early detection and intervention for mental health among users of social media.