Speech Emotion Recognition and Personalized Mental Health Recommendations Using DSCNN and BERT
摘要
In recent years, there has been a rise in the prevalence of mental health disorders, which has highlighted the urgent need for readily available, non-invasive diagnostic techniques. This study offers a novel framework for assessing mental health by fusing a Deep Stride Convolutional Neural Network (DSCNN) for recognition of emotion with BERT (Bidirectional Encoder Representations from Transformers) for contextual analysis and recommendation creation. The DSCNN is engineered to directly extract high-level emotional aspects from unprocessed speech data by efficiently capturing temporal correlations and subtle changes in acoustic parameters through its stride-based architecture. Subsequently, the extracted features are input into a refined BERT model trained on the combination of emotion-labeled datasets and text corpora related to mental health to facilitate accurate sentiment analysis and complex contextual understanding. The proposed system does not only identify the emotional state but also generates precautions according to an individual's emotional state by correlating detected emotions with specific mental health conditions. Delivering real-time emotional insights and personalized support thus bridges the gap between the pace of technological advancement and the compassion inherent in the management of mental health, empowering self-control over one's emotional well-being.