A Natural Language Processing Approach for Early Detection of Suicide and Depression Risk on Social Media Platforms
摘要
This paper presents an NLP-based approach for the early detection of mental health risks, focusing on depression and suicidal ideation. The study uses publicly available datasets from Reddit, Kaggle, and Mendeley, applying a dual scoring strategy (cleaned vs. original text) to enhance interpretability. The model demonstrates solid performance: when classifying texts into two categories (risk/no risk), it achieves an accuracy of 86.02% with an F1-score of 0.84 and recall of 0.85. When expanded into three categories (depression/suicide/none), accuracy decreases to 71.09%, with an average F1-score of 0.70. These results highlight the challenge of differentiating closely related emotional states while ensuring explainability. Beyond technical contributions, the study emphasizes the ethical use of NLP in mental health, underlining that such models should serve as supportive tools and not replacements for professional intervention.