错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Gaussian Distribution Labeling Method for Speech Quality Assessment

  • Minh Tu Le,
  • Bao Thang Ta,
  • Nhat Minh Le,
  • Phi Le Nguyen,
  • Van Hai Do

摘要

Speech Quality Assessment (SQA) plays a critical role in speech processing systems by predicting the quality of input speech signals. Traditionally, quality is represented as mean of scores provided by many listeners’ opinions and is a floating-point value, so the straightforward approach is treating SQA as a regression task. However, speech quality score is labeled by crowdsourcing, resulting in a lack of uniformity among listeners. This discrepancy makes it challenging to accurately represent speech quality. Some recent works have proposed transforming SQA into a classification task by converting continuous scores into discrete classes. These approaches simplify the training of SQA models but not consider valuable variance information, which capture the differences among listeners and the difficulties to evaluate. To address this limitation, in addition to directly learning score values through the regression task, we propose using a classification task in a multitask setting. We assume that quality labels for each class follow a Gaussian distribution, considering not only the mean ratings but also the variance among them. To ensure that the improvements can be confidently attributed to our approach, we adopt a Conformer-based architecture, a state-of-the-art architectural framework for representing speech information, as the core for our model and baseline models. Through a comprehensive series of experiments conducted on the NISQA speech quality assessment dataset, we provide conclusive evidence that our proposed method consistently enhances results across all dimensions of speech quality.