Investigating Automated Transcriptions for Multimodal CPS Detection in Groupwork
摘要
Monitoring collaborative learning in educational settings is challenging, because teachers cannot track all student groups in detail. Recently, numerous works have shown that collaborative learning can be automatically tracked with machine learning. In this work, we explore how Google and Whispers automatic transcription systems influence Collaborative Problem Solving (CPS) classification performance using a public dataset. We also look at how different styles of transcribing overlapping speech influence CPS classification performance. The transcription styles in our analysis are active speaker (transcriptions of a single active speaker) and temporal (transcriptions of all participants in speaking order). As a baseline, manually segmented speech utterances were transcribed for each transcription style. We first evaluated Automatic Speech Recognition (ASR) performance as it pertains to each type. Then, we evaluate how each type influences CPS classification. We find that active speaker transcription equates with higher CPS classification performance, but that Google ASR is closer to temporal transcriptions. Our results provide insights into how to capitalize on ASRs in collaborative contexts.