<p>Identifying​‍​‌‍​‍‌​‍​‌‍​‍‌ mental health problems at an early stage should be a top priority if one is to have timely intervention. This research introduces a multimodal system that combines text, audio, and video data by applying deep learning and generative AI ​‍​‌‍​‍‌​‍​‌‍​‍‌techniques. Missing or incomplete data are addressed through Multiple Imputation with Graph-Based Filtering (MI-GBF), preserving data integrity. Every​‍​‌‍​‍‌​‍​‌‍​‍‌ modality is first analyzed on its own and then combined through the Fusion-Three Branch Network (F-TBN) to obtain additional information. To improve natural language understanding, BERT (Bidirectional Encoder Representations from Transformers) is used, which helps in context-aware evaluation. Such a method serves as a scalable, flexible instrument for the initial identification and the personalized clinical evaluation of mental health disorders; hence, it is a nice showcase of the generative AI potential in the multimodal healthcare ​‍​‌‍​‍‌​‍​‌‍​‍‌sector. Experimental results indicate high predictive performance, with video-based facial expression analysis achieving 90% accuracy and speech recordings 88%, demonstrating the framework’s effectiveness in capturing emotional cues. This approach can be extended to real-time monitoring and adaptive intervention systems for mental health care.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Early detection of mental health conditions using multimodal generative AI with MI-GBF and fusion-three branch network

  • Amit Kumar Saxena,
  • Rajeshwar Prasad,
  • Suman Laha

摘要

Identifying​‍​‌‍​‍‌​‍​‌‍​‍‌ mental health problems at an early stage should be a top priority if one is to have timely intervention. This research introduces a multimodal system that combines text, audio, and video data by applying deep learning and generative AI ​‍​‌‍​‍‌​‍​‌‍​‍‌techniques. Missing or incomplete data are addressed through Multiple Imputation with Graph-Based Filtering (MI-GBF), preserving data integrity. Every​‍​‌‍​‍‌​‍​‌‍​‍‌ modality is first analyzed on its own and then combined through the Fusion-Three Branch Network (F-TBN) to obtain additional information. To improve natural language understanding, BERT (Bidirectional Encoder Representations from Transformers) is used, which helps in context-aware evaluation. Such a method serves as a scalable, flexible instrument for the initial identification and the personalized clinical evaluation of mental health disorders; hence, it is a nice showcase of the generative AI potential in the multimodal healthcare ​‍​‌‍​‍‌​‍​‌‍​‍‌sector. Experimental results indicate high predictive performance, with video-based facial expression analysis achieving 90% accuracy and speech recordings 88%, demonstrating the framework’s effectiveness in capturing emotional cues. This approach can be extended to real-time monitoring and adaptive intervention systems for mental health care.