Recent dialogues increasingly include not only text but also images. However, research on multimodal dialogue embeddings remains limited, and existing methods face three primary challenges: i) they often reflect only the context of specific moments when images are shared, failing to capture the overall meaning of the dialogue, ii) they require large batch sizes to achieve high performance when using contrastive learning, and iii) they generally lack evaluation of intrinsic tasks that directly measure the structural quality and semantic consistency of embeddings, as they rely heavily on extrinsic tasks. To address these issues, we propose MMCDE, multimodal contrastive learning for dialogue embeddings with global and local views. Our method constructs contrastive pairs by leveraging a global view that considers the context of the entire dialogue and a local view that captures interactions between images and text within the dialogue. We are the first to include both extrinsic and intrinsic tasks in the evaluation of performance in multimodal dialogue embedding research. Furthermore, we demonstrate that our approach achieves superior performance across three tasks, even with limited memory and batch sizes. Our code is available at https://github.com/subeenc/MMCDE our GitHub repository.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Contrastive Learning for Dialogue Embeddings with Global and Local Views

  • Subeen Choe,
  • Jihyeon Oh,
  • Jihoon Yang

摘要

Recent dialogues increasingly include not only text but also images. However, research on multimodal dialogue embeddings remains limited, and existing methods face three primary challenges: i) they often reflect only the context of specific moments when images are shared, failing to capture the overall meaning of the dialogue, ii) they require large batch sizes to achieve high performance when using contrastive learning, and iii) they generally lack evaluation of intrinsic tasks that directly measure the structural quality and semantic consistency of embeddings, as they rely heavily on extrinsic tasks. To address these issues, we propose MMCDE, multimodal contrastive learning for dialogue embeddings with global and local views. Our method constructs contrastive pairs by leveraging a global view that considers the context of the entire dialogue and a local view that captures interactions between images and text within the dialogue. We are the first to include both extrinsic and intrinsic tasks in the evaluation of performance in multimodal dialogue embedding research. Furthermore, we demonstrate that our approach achieves superior performance across three tasks, even with limited memory and batch sizes. Our code is available at https://github.com/subeenc/MMCDE our GitHub repository.