<p>To address the challenge encountered in audio–visual multimodal learning where heterogeneous learning rates result in one modality dominating the learning process and suppressing other modalities, thus weakening the multimodal collaborative decision-making process, a multimodal adaptive balanced learning method based on gradient modulation (AGM-CR) is proposed. This method introduces modulation coefficients to dynamically adjust the learning rates of different modalities based on gradient variations. A gradient balancing strategy is further employed by incorporating the gradient losses of individual modalities into the total loss as a regularization term to mitigate gradient disparities and balance the learning process. Experimental results show that AGM-CR improves classification accuracy by 3.1% and 1.3%&#xa0;on the CREMA-D and RAVDESS datasets, respectively, and reduces gradient fluctuations over multiple iterations, thereby enhancing training stability and accelerating convergence. Moreover, AGM-CR is a plug-and-play approach, offering greater flexibility and generalizability than the existing balancing methods do.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An audio–visual multimodal adaptive balanced learning method based on gradient modulation

  • Wenxiu Ao,
  • Zhongmei Wang,
  • Jianhua Liu,
  • Shenao Peng,
  • Liang Zheng

摘要

To address the challenge encountered in audio–visual multimodal learning where heterogeneous learning rates result in one modality dominating the learning process and suppressing other modalities, thus weakening the multimodal collaborative decision-making process, a multimodal adaptive balanced learning method based on gradient modulation (AGM-CR) is proposed. This method introduces modulation coefficients to dynamically adjust the learning rates of different modalities based on gradient variations. A gradient balancing strategy is further employed by incorporating the gradient losses of individual modalities into the total loss as a regularization term to mitigate gradient disparities and balance the learning process. Experimental results show that AGM-CR improves classification accuracy by 3.1% and 1.3% on the CREMA-D and RAVDESS datasets, respectively, and reduces gradient fluctuations over multiple iterations, thereby enhancing training stability and accelerating convergence. Moreover, AGM-CR is a plug-and-play approach, offering greater flexibility and generalizability than the existing balancing methods do.