Detection of Depression in Spanish Texts Using BERT-Based Language Model
摘要
The Mexican Ministry of Health collects mental health data for medical resource allocation, highlighting depression with a 5.3% prevalence. Cultural biases, socioeconomic factors and the lack of specialists affect its diagnosis and treatment. To improve early detection, the use of BERT-based models trained to detect depression on Spanish-written texts describing the condition, thanks to its technology for understanding and classifying contexts. Training process using the Depression dataset from Kaggle and cross-validation with k = 5 was applied to prevent overfitting. Testing and validation used accuracy, precision, recall, and F1-score metrics, demonstrating that BERT-based models outperform traditional machine learning techniques in detecting depression. The trained model was applied to a survey collected on traumatic events experienced by university-level students.