A Pilot Study on the Prosodic Factors Influencing Voice Attractiveness of AI Speech
摘要
This study investigates the attractiveness of AI-synthesized voice and its influencing factors from the perspective of speech prosody. Firstly, a comparative analysis, including MOS (Mean Opinion Score) listening test, was conducted on the vocal and facial attractiveness levels of four most popular AI-synthesized voices with 20 listeners, and the best-performing ChatGPT (Chat Generative Pre-trained Transformer) was selected for subsequent analysis. Secondly, a segment of ChatGPT speech was re-synthesized using Praat and Python with different prosodic acoustic parameters, such as overall fundamental frequency, intonation, and duration, respectively. Then a MOS listening test was carried out to evaluate the voice attractiveness of re-synthesized segments in four dimensions: power, competence, warmth, and honesty. The statistical results, including ANOVA (Analysis of Variance) and HSD (Honestly Significant Difference), indicate that altering prosodic parameters, particularly an overly high fundamental frequency and an excessively high speech rate, would significantly reduce the attractiveness of AI-synthesized voice, while the opposite is not the same. Thirdly, to exclude semantic factors, a speech segment in Finnish, which is a completely unfamiliar language to listeners, was used for MOS listening test, and the results are consistent in different languages.