Tracking Individual Beliefs in Co-situated Groups Using Multimodal Input
摘要
In co-situated collaborative groups, a challenge for automated interpretation of group dynamics is parsing and attributing input from individual group members to process their respective perspectives and contributions. In this work, we describe the necessary components for such a system to handle multimodal, multi-party input. We apply these methods over an audiovisual dataset of a co-situated collaborative task called the Weights Task Dataset (WTD) to track individual beliefs regarding the task. We find that combining audiovisual speaker detection (ASD) with utterance transcripts enables us to track individuals’ beliefs during a task. We show that our system succeeds in individual belief tracking, achieving scores similar to those seen in dense-paraphrased common ground tracking. Further, we demonstrate that a combination of ASD and point target detection can be applied to transcripts for automated dense paraphrasing. We additionally identify where individual components need to be improved, including ASD and task-belief identification.