AI has recently started to gain popularity in most spheres of our daily life, including healthcare. While most AI-based solutions have been proposed for physical-related medical conditions, only few of them target mental health. AI-based solutions can support the diagnosis of mental disorders from clues that a human doctor may not be able to recognize, for example the biomarkers of depression. Previous studies in the area of automatic detection of depression from different media, like audio-visual data, ECG, or transcribed speech exists. In this research, we focus on the detection of depression from audio data. We use only low-level audio features and their physical characteristics, while we do not include any semantic information carried by the speech. We performed a qualitative analysis on the selected set of features that proved to be important for recognizing depression. Such an analysis revealed the dependence between gender and the set of relevant features for depression detection. We discovered differences and similarities between sets of strong predictors between gender. Formants have shown to be important to describe the articulation level for both male and female voices. Control over the voice is a better predictor for male voices, while monotony is better described with formants for male voices and with MFCC and energy for female ones. This highlights the importance of gender inclusivity in mental healthcare and depression diagnostics in particular. Experiments carried out while running this study achieved a value of F1 of 65% for depression detection for female voices and 67% for male voices.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detecting Depression from Audio Data

  • Mary Idamkina,
  • Andrea Corradini

摘要

AI has recently started to gain popularity in most spheres of our daily life, including healthcare. While most AI-based solutions have been proposed for physical-related medical conditions, only few of them target mental health. AI-based solutions can support the diagnosis of mental disorders from clues that a human doctor may not be able to recognize, for example the biomarkers of depression. Previous studies in the area of automatic detection of depression from different media, like audio-visual data, ECG, or transcribed speech exists. In this research, we focus on the detection of depression from audio data. We use only low-level audio features and their physical characteristics, while we do not include any semantic information carried by the speech. We performed a qualitative analysis on the selected set of features that proved to be important for recognizing depression. Such an analysis revealed the dependence between gender and the set of relevant features for depression detection. We discovered differences and similarities between sets of strong predictors between gender. Formants have shown to be important to describe the articulation level for both male and female voices. Control over the voice is a better predictor for male voices, while monotony is better described with formants for male voices and with MFCC and energy for female ones. This highlights the importance of gender inclusivity in mental healthcare and depression diagnostics in particular. Experiments carried out while running this study achieved a value of F1 of 65% for depression detection for female voices and 67% for male voices.