错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Learning Algorithm for Image Classification Integrating Multimodal Information

  • Zhenwei Di,
  • Xiaotong Huang

摘要

Objective: Although multimodal data has been used extensively, single model recognition represents the majority of current recognition technologies, which has a poor recognition effect in complicated contexts. This research presents a multimodal fusion approach technique combining convolutional neural network (CNN) and self-attention mechanism to enhance the accuracy of picture categorization under multimodal input. Methods: Five categories—cats, dogs, humans, automobiles, and flowers—as well as 10,000 training sets, 5,000 validation sets, and 5,000 test sets are employed in the COCO multimodal dataset. The self-attention mechanism is used to weightedly fuse the features of each modality, improve the correlation between multi-modalities, and increase the classification accuracy of images after the convolutional neural network method has been used to extract features of images of various modalities in order to obtain feature information. Results: The experimental results show that the accuracy of a single model image is 85.2%, while the accuracy of the model after multimodal fusion is 90.1%. Therefore, image classification using fused multimodal information is better than traditional single modal data. The fusion of multimodal information significantly improves the accuracy of the image model. Conclusion: By using convolutional neural networks and self-attention mechanisms to fuse multimodal information for image classification, the accuracy of information images is greatly improved.