错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EFCC-IeT: Cross-Modal Electronic File Content Correlation via Image-Enhanced Text

  • Pengfei Jing,
  • Jiguo Liu,
  • Chao Liu,
  • Meimei Li

摘要

With the development of the information age and the popularization of electronic documents, a large number of electronic documents are generated every day in government agencies, enterprises and institutions, and are used more and more frequently. Domain experts need to analyze the contents of these documents to detect whether there is sensitive information. However, the current traditional electronic document content inspection is only limited to a single mode, ignoring the inspection method that combines image and text features. Aiming at the problem of missing modalities for electronic document content inspection, we propose a Cross-modal Electronic File Content Correlation method via Image-enhanced Text (EFCC-IeT), which extracts text and image features respectively. Then we designed a combined attention image-text feature fusion module based on the Transformer encoder and the attention mechanism. By calculating the information correlation between each word in the text and the image, the representation ability of text features is improved. Finally, the text features and image features calculated by combined attention are concatenated and input into the fully connected layer. Experimental results show that the precision, recall rate and F1 value of the EFCC-IeT method reached 82.09%, 86.50% and 74.63% respectively on the Conceptual Captions dataset. Compared with the single-modal method and other benchmark models, it has greater performance improvement, which verifies the effectiveness of the method.