With the development of information technology, multimodal aspect level analysis tasks have become a hot research topic. By fully mining the features of each modality, alignment and fusion of each modality can be achieved, promoting information complementarity between modalities and achieving better classification results. Therefore, an introduction can be provided from three aspects: multimodal feature extraction, aspect level analysis, and multimodal contrastive learning. Previous studies have shown that English vocabulary learning strategies in multimodal environments have not received sufficient attention. The existing research mainly focuses on single mode learning, fails to fully consider the learners’ ability to acquire vocabulary through visual and contextual information, and fails to completely solve the dilemma of English vocabulary learning in a multimodal environment. This article aims to propose and validate a multimodal learning strategy based on VGG16 (Visual Geometry Group Network-16) and BERT (Bidirectional Encoder Representation from Transformers) models to optimize the effectiveness of English vocabulary learning in multimodal environments. By integrating visual and contextual information, this article utilizes deep learning models to extract features and uses corresponding learning algorithms for training, in order to achieve better learning results and application prospects. Traditional English vocabulary learning has limitations in a single modal environment, which can not make full use of visual and contextual information. This paper proposes a multimodal learning strategy based on VGG16 and BERT model to optimize the learning effect. The experimental results demonstrate that the model achieves the highest accuracy when the learning rate is 0.001 and the batch size is 64.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Application of Data Mining Based on BERT and Decision Tree in English Vocabulary Analysis

  • Shuang Zhang,
  • Yu Han

摘要

With the development of information technology, multimodal aspect level analysis tasks have become a hot research topic. By fully mining the features of each modality, alignment and fusion of each modality can be achieved, promoting information complementarity between modalities and achieving better classification results. Therefore, an introduction can be provided from three aspects: multimodal feature extraction, aspect level analysis, and multimodal contrastive learning. Previous studies have shown that English vocabulary learning strategies in multimodal environments have not received sufficient attention. The existing research mainly focuses on single mode learning, fails to fully consider the learners’ ability to acquire vocabulary through visual and contextual information, and fails to completely solve the dilemma of English vocabulary learning in a multimodal environment. This article aims to propose and validate a multimodal learning strategy based on VGG16 (Visual Geometry Group Network-16) and BERT (Bidirectional Encoder Representation from Transformers) models to optimize the effectiveness of English vocabulary learning in multimodal environments. By integrating visual and contextual information, this article utilizes deep learning models to extract features and uses corresponding learning algorithms for training, in order to achieve better learning results and application prospects. Traditional English vocabulary learning has limitations in a single modal environment, which can not make full use of visual and contextual information. This paper proposes a multimodal learning strategy based on VGG16 and BERT model to optimize the learning effect. The experimental results demonstrate that the model achieves the highest accuracy when the learning rate is 0.001 and the batch size is 64.