错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Tri-factorized Modular Hypergraph Autoencoder for Multimodal Semantic Analysis

  • Shaily Malik,
  • Geetika Dhand,
  • Kavita Sheoran,
  • Divya Jatain,
  • Vaani Garg

摘要

For image-to-text and text-to-image classifications, the features of data collected from various imaging devices, sensors, and their text descriptions must be mapped into a common latent space with reduced dimensions. The low-dimensional features are supposed to provide the most information with the least amount of loss.In this paper we propose a cross-modal semantic autoencoder that uses nonnegative matrix factorization (NMF) to factorize the features into a lower rank. Due to two matrix factorization, the traditional NMF is unable to translate all of the information into lower space. This is addressed by a unique tri-factorized NMF with hypergraph regularization. Instead of using the feature adjacency matrix in hypergraph regularization, a more information-rich modularity matrix is suggested. The Wiki dataset is used to evaluate this tri-factorized hypergraph regularized multimodal autoencoder for image-to-text and text-to-image conversion. In order to lower the feature dimension, Multimodal Conditional Principal label space transformation (MCPLST) is also enabled by this novel autoencoder. Comparing the proposed autoencoder against the semantic autoencoder, the former showed an improvement in classification accuracy of up to 1.8%.