Exploring Multimodal Information Fusion in Spoken Off-Topic Degree Assessment
摘要
Currently, most research methods for spoken off-topic detection are based on the results of upstream speech recognition tasks. However, upstream speech recognition tasks may introduce issues such as homophones, text recognition errors, and semantic confusion. Therefore, this study aims to explore the impact of integrating audio features into text information on off-topic degree assessment. To achieve this goal, we collected a dataset consisting of 2652 responses in question-answering scenarios from public competitions. The data was annotated according to the evaluation guidelines for farmers and herdsmen, creating a dataset named ASAG-TD for assessing off-topic degree in question-answering scenarios. In addition, we conducted research using the pre-trained language model RoBERTa, combined with commonly used neural network models, exploring two aspects: audio features and self-supervised pre-trained acoustic model. Experimental results demonstrate the effectiveness of our method, with the mean absolute error (MAE) of off-topic degree in spoken responses reduced to 0.414 and a Pearson correlation coefficient of 0.95.