Prediction of self-reported and external observations of conversational engagement in online group discussions
摘要
Automated engagement detection has to date focussed exclusively on engagement annotations collected via external observation. It is therefore unknown whether model predictions align with a participant’s own perception of their engagement level. In this study, we examine self-reported and external observations of engagement in a corpus of small-group conversational interactions. We find no evidence of correlation between self-reported and externally observed engagement. We show that annotators rely heavily on speaking activity as a proxy for engagement, which we demonstrate is a poor indicator of self-reported engagement. The focus of recent literature has been to improve the engagement detection accuracy with deep learning. In contrast, we use a simple multimodal combination of low-level linguistic features (e.g. number of words spoken) and facial expression. We find that none of the linguistic features significantly predict self-reported engagement. However, we find that each linguistic feature significantly predicts annotator engagement, with the odds of a high engagement prediction increasing with each word and utterance spoken. Finally, we report the performance of a model which predicts real-time engagement levels. We find that the best model overall utilises a multimodal combination of linguistic and facial expression features. Our work raises a number of important questions concerning the ecological validity of current approaches for engagement detection and underlines the need to validate external observations against self-reported metrics.