A hierarchical multimodal intelligent learning framework for cognitive diagnosis and learning performance prediction
摘要
Based on the existing paradox between the need for precise cognitive diagnosis and the insufficiency of single-modal evaluation in intelligent learning systems, a hierarchical multimodal fusion architecture is proposed. This architecture adopts the paradigm of “encoding-fusion-diagnosis”, and it combines the decoding-enhanced bidirectional encoder representation from transformers with disentangled attention (DeBERTa), contrastive language-image pretraining (CLIP)-vision transformer (ViT), and wave to vector (Wav2Vec) 2.0 encoders. Experimental performance indicates that the cognitive diagnosis (Micro-F1 = 0.78) and the prediction of problem difficulty (root mean squared error (RMSE) = 0.126) perform greatly improved compared to the baseline model. It provides a major benefit concerning more precise cognitive diagnosis ability for complicated reasoning tasks, adaptive allocation of the attention ratio of each modality according to the type of the questioned task, superb early prediction performance, and superior robustness on noisy text. This study illustrates that multimodal fusion technology is extremely efficient for the deep cognitive state evaluation task. It provides a major technical route for the construction of a brand-new intelligent learning system, and the proposed hierarchical multimodal fusion framework can serve as the core cognitive assessment module.