ATFN: An Efficient Multi-modal Depression Assistance Diagnostic Model Based on Multi-channel Attention Mechanism
摘要
The vast population of depression patients and the inefficiency in their identification pose a significant challenge in the current medical field. In light of this, this paper proposes a method that integrates the wisdom of manual feature design with the advantages of deep learning models, constructing the ATFN, a multi-modal fusion depression assistance diagnostic model based on a multi-channel attention mechanism. Validated on the public depression dataset E-DAIC, the results show that the audio-modality-based auxiliary diagnostic model (Audio) achieves an accuracy of 90.6%, an F1 score of 0.88, and simultaneously exhibits a sensitivity of 84.6% and a specificity of 90%. The text-modality-based auxiliary diagnostic model (Text), while having a slightly lower specificity of 75%, achieves a recall rate and sensitivity of 100%, with an F1 score of 0.83. When we fuse the multi-modal features, the diagnostic performance of the model is significantly improved, with the accuracy rising to 91.7%, the F1 score reaching 0.92, and maintaining a sensitivity of 100% and a specificity of 85.7%. These results validate the effectiveness and superiority of the proposed method in this paper.