<p>Automated engagement detection has to date focussed exclusively on engagement annotations collected via external observation. It is therefore unknown whether model predictions align with a participant’s own perception of their engagement level. In this study, we examine self-reported and external observations of engagement in a corpus of small-group conversational interactions. We find no evidence of correlation between self-reported and externally observed engagement. We show that annotators rely heavily on speaking activity as a proxy for engagement, which we demonstrate is a poor indicator of self-reported engagement. The focus of recent literature has been to improve the engagement detection accuracy with deep learning. In contrast, we use a simple multimodal combination of low-level linguistic features (e.g. number of words spoken) and facial expression. We find that none of the linguistic features significantly predict self-reported engagement. However, we find that each linguistic feature significantly predicts annotator engagement, with the odds of a high engagement prediction increasing with each word and utterance spoken. Finally, we report the performance of a model which predicts real-time engagement levels. We find that the best model overall utilises a multimodal combination of linguistic and facial expression features. Our work raises a number of important questions concerning the ecological validity of current approaches for engagement detection and underlines the need to validate external observations against self-reported metrics.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prediction of self-reported and external observations of conversational engagement in online group discussions

  • Sam O’Connor Russell,
  • Justine Reverdy,
  • Benjamin Cowan,
  • Naomi Harte

摘要

Automated engagement detection has to date focussed exclusively on engagement annotations collected via external observation. It is therefore unknown whether model predictions align with a participant’s own perception of their engagement level. In this study, we examine self-reported and external observations of engagement in a corpus of small-group conversational interactions. We find no evidence of correlation between self-reported and externally observed engagement. We show that annotators rely heavily on speaking activity as a proxy for engagement, which we demonstrate is a poor indicator of self-reported engagement. The focus of recent literature has been to improve the engagement detection accuracy with deep learning. In contrast, we use a simple multimodal combination of low-level linguistic features (e.g. number of words spoken) and facial expression. We find that none of the linguistic features significantly predict self-reported engagement. However, we find that each linguistic feature significantly predicts annotator engagement, with the odds of a high engagement prediction increasing with each word and utterance spoken. Finally, we report the performance of a model which predicts real-time engagement levels. We find that the best model overall utilises a multimodal combination of linguistic and facial expression features. Our work raises a number of important questions concerning the ecological validity of current approaches for engagement detection and underlines the need to validate external observations against self-reported metrics.