As a broader range of natural language processing tools become available, from large language models to more natural-sounding text-to-speech resources, new questions emerge about how to curate robot voices. One tempting option may be to use by default the latest tool or most human-sounding voice option. But is this the right path, or in fact a misstep? This question, and a broader wondering about whether and how adjusting the humanness of the voice used by our lab’s robotic stand-up comedian would matter, led to the design of the presented experiment, which evaluated the effects of three different levels of voice humanness on peoples’ perceptions of our robotic comedian. This online and between-subjects video-based study ( \(N = 91\) ) showed that both the least and the most human-like voices caused detriments in ratings of the robot on a subset of the evaluation scales, without offering any discernible benefit to the interaction. Based on these observations, we urge the robotics community to select voices with care and consider that matching the chosen voice to the presentation of the robot itself will likely lead to a maximally successful interaction experience.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dude, Where’s My Robot Voice? Sometimes More Robotic Is Better in Social Robot Speech Generation

  • Christopher A. Sanchez,
  • Timothy Bui,
  • Naomi T. Fitter

摘要

As a broader range of natural language processing tools become available, from large language models to more natural-sounding text-to-speech resources, new questions emerge about how to curate robot voices. One tempting option may be to use by default the latest tool or most human-sounding voice option. But is this the right path, or in fact a misstep? This question, and a broader wondering about whether and how adjusting the humanness of the voice used by our lab’s robotic stand-up comedian would matter, led to the design of the presented experiment, which evaluated the effects of three different levels of voice humanness on peoples’ perceptions of our robotic comedian. This online and between-subjects video-based study ( \(N = 91\) ) showed that both the least and the most human-like voices caused detriments in ratings of the robot on a subset of the evaluation scales, without offering any discernible benefit to the interaction. Based on these observations, we urge the robotics community to select voices with care and consider that matching the chosen voice to the presentation of the robot itself will likely lead to a maximally successful interaction experience.