The end-to-end training of neural networks with multimodal data poses challenges beyond those observed for the training with unimodal data. The difficulty lies frequently in a network’s capacity to overfit and generalize at different rates in each modality, resulting networks may thus perform worse than networks trained on unimodal data alone. In the present study, we introduce a late fusion methodology designed for the bimodal binary classification task under the assumption that the logit spaces of both modalities can be modeled by a joint Gaussian distribution. On the one hand, it is suitable when ample bimodal data is available, avoiding the effect of catastrophic fusion, where unimodal models may perform better than the bimodal systems. On the other hand, the proposed technique is particularly well suited for situations in which significant amounts of unimodal data are available to train separate networks for each modality, but where only a limited amount of joint, bimodal data is available to train the fusion mechanism. The proposed method is assessed over the PAN 2018 gender identification dataset, which involves the identification of an authors’ (binary) gender label through their tweets and posted images. The cross-validation results demonstrate that our proposed hybrid model-based method exhibits superior performance and robustness compared to competing state-of-the-art neural-network-based approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mitigating Overfitting in Bimodal Data Integration with GLOW for Large and Sparse Datasets

  • Wentao Yu,
  • Dorothea Kolossa,
  • Robert Nickel

摘要

The end-to-end training of neural networks with multimodal data poses challenges beyond those observed for the training with unimodal data. The difficulty lies frequently in a network’s capacity to overfit and generalize at different rates in each modality, resulting networks may thus perform worse than networks trained on unimodal data alone. In the present study, we introduce a late fusion methodology designed for the bimodal binary classification task under the assumption that the logit spaces of both modalities can be modeled by a joint Gaussian distribution. On the one hand, it is suitable when ample bimodal data is available, avoiding the effect of catastrophic fusion, where unimodal models may perform better than the bimodal systems. On the other hand, the proposed technique is particularly well suited for situations in which significant amounts of unimodal data are available to train separate networks for each modality, but where only a limited amount of joint, bimodal data is available to train the fusion mechanism. The proposed method is assessed over the PAN 2018 gender identification dataset, which involves the identification of an authors’ (binary) gender label through their tweets and posted images. The cross-validation results demonstrate that our proposed hybrid model-based method exhibits superior performance and robustness compared to competing state-of-the-art neural-network-based approaches.