Application of Deep Learning Models for Multimodal Information Fusion in Intelligent Information Classification
摘要
With the rapid growth of multimedia information, the need for more effective classification techniques has become crucial. In this paper, we propose a deep learning model that utilizes a mechanism for fusing multimodal information to perform fine-grained image classification in intelligent information. The method aims to leverage both visual and textual semantic information for joint reasoning of image categories. To extract rich and comprehensive information from images, the algorithm extracts three different basic information types from images through a global image information extraction module, a local object information extraction module, and a local text information extraction module. To explore and utilize the potential connections between different modalities of information, the algorithm fuses the three basic information types through an information interaction fusion module, ultimately using enhanced image representation features for classification prediction. By combining textual and visual information, the model improves classification accuracy compared to single-modal methods. We present the architecture of this model and demonstrate its effectiveness through extensive experiments. Our research highlights the potential of multimodal fusion in enhancing the process of intelligent information classification.