GSIP: A New System for Prosody Selection for Gibberish Speech
摘要
We present the development and evaluation of GSIP (Gibberish Speech Impression Predictor), a bi-directional GRU neural network which serves as a human impression prediction model for speech inputs, incorporating both phonetic information and prosody matrices. The objective is to select appropriate acoustic prosody for gibberish and semantic speech. Experimental validation of the proposed system was conducted through a user study. Participants ranked the system’s performance against constant prosody and random prosody patterns and completed adapted Godspeed scale questionnaires to assess their perception of the GSIP-based prosody system and conversational agents. The experiment employed three embodied conversational agents, two screen-based avatars and a physical robot. The results suggest: 1) Gibberish Speech is not so engaging for conversation; 2) higher anthropomorphism degrees create a higher perception of intelligence when agents spoke Gibberish, but that effect did not hold when speaking English. We also found that our proposed system accurately predicts human impression, but fails to generate more engaging reactions compared to constant prosody.