Speech quality plays a crucial role in the audio transcription process as it influences various factors that could significantly impact the annotation results. These factors include transcription speed, annotation confidence, and the number of playbacks, among others. Most existing subjective (e.g., Mean Opinion Score (MOS)) and objective (e.g., Perceptual Evaluation Score Quality (PESQ)) speech quality measurements do not consider factors that could hinder the transcription process. In this work, we first analyze the relationship between subjective speech quality measurement MOS and factors that may impact the transcription quality. Findings show that this measure poorly correlates with the speech quality perceived by the annotator. Then, we explore the use of a novel subjective measurement, called Speech Quality Score (SQS), whose scale aims to encompass the most relevant factors involved in the audio transcription process. Additionally, we explore two different approaches to predict the SQS measure. The results indicate a linear correlation of ( \(r=0.66\) ) and ( \(r=0.34\) ) between the ground truth and the predicted SQS values on the VoxConverse and RTVE2020 datasets, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

How Does Speech Quality Impact the Data Transcription Process?

  • Fernando M. Espinoza-Cuadros,
  • Rafael Ginard-Aguilera,
  • Juan M. Perero-Codosero

摘要

Speech quality plays a crucial role in the audio transcription process as it influences various factors that could significantly impact the annotation results. These factors include transcription speed, annotation confidence, and the number of playbacks, among others. Most existing subjective (e.g., Mean Opinion Score (MOS)) and objective (e.g., Perceptual Evaluation Score Quality (PESQ)) speech quality measurements do not consider factors that could hinder the transcription process. In this work, we first analyze the relationship between subjective speech quality measurement MOS and factors that may impact the transcription quality. Findings show that this measure poorly correlates with the speech quality perceived by the annotator. Then, we explore the use of a novel subjective measurement, called Speech Quality Score (SQS), whose scale aims to encompass the most relevant factors involved in the audio transcription process. Additionally, we explore two different approaches to predict the SQS measure. The results indicate a linear correlation of ( \(r=0.66\) ) and ( \(r=0.34\) ) between the ground truth and the predicted SQS values on the VoxConverse and RTVE2020 datasets, respectively.