<p>This study aims to evaluate the quality of mimicked speech by integrating spectral and prosodic speaker embeddings to identify the best mimicked artist, assessed via Mean Opinion Score testing. Augmented d-vector embeddings are constructed separately from prosodic features (loudness, pitch, speaking rate, shimmer, tempogram ratio) and spectral features (Mel Frequency Cepstral Coefficients, chroma, tonnetz, spectral roll off, bandwidth, flux, centroid). Each embedding type is processed and classified independently using a Deep Neural Network classifier. The classifiers produce prediction probability scores for all candidate artists, which are then fused at the score level using empirically chosen weighting constants to generate a combined ranking and select the top-1 predicted mimicked artist. Experiments conducted on the MIMICz dataset achieve 68% top-1 accuracy, outperforming early fusion and baseline methods, while preserving the distinct representational contributions of both spectral and prosodic feature sets. To the best of our knowledge, this is the first work to apply score level fusion of augmented spectral and prosodic embeddings for speech mimicry recognition, combining the strengths of both feature domains without compromising their individual contributions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Learning-Based Assessment of Impersonated Speech Using Score-Level Fusion of Augmented D-Vector Embeddings

  • Bhasi K.C.,
  • Rajeev Rajan

摘要

This study aims to evaluate the quality of mimicked speech by integrating spectral and prosodic speaker embeddings to identify the best mimicked artist, assessed via Mean Opinion Score testing. Augmented d-vector embeddings are constructed separately from prosodic features (loudness, pitch, speaking rate, shimmer, tempogram ratio) and spectral features (Mel Frequency Cepstral Coefficients, chroma, tonnetz, spectral roll off, bandwidth, flux, centroid). Each embedding type is processed and classified independently using a Deep Neural Network classifier. The classifiers produce prediction probability scores for all candidate artists, which are then fused at the score level using empirically chosen weighting constants to generate a combined ranking and select the top-1 predicted mimicked artist. Experiments conducted on the MIMICz dataset achieve 68% top-1 accuracy, outperforming early fusion and baseline methods, while preserving the distinct representational contributions of both spectral and prosodic feature sets. To the best of our knowledge, this is the first work to apply score level fusion of augmented spectral and prosodic embeddings for speech mimicry recognition, combining the strengths of both feature domains without compromising their individual contributions.