错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adapting Audiovisual Speech Synthesis to Estonian

  • Sven Aller,
  • Mark Fishel

摘要

Audiovisual speech synthesis is an important topic from different points of view: visualization helps to understand speech more easily in noisy environments, conversation is more natural, and clarity is much better for hearing-impaired as well as other users. At the same time, its availability is limited to a much narrower selection of languages than speech-only synthesis, and language-independent methods of adding the visual part are not thoroughly tested for most languages. This paper presents the development of two methods of adapting audiovisual speech synthesis to Estonian. We reuse an existing neural speech synthesis model and adapt a speech-driven and text-driven approach to adding the visual part. We contrast the two developed solutions with pure audio in conditions with different noise levels and evaluate the clarity, naturalness, and pleasantness of the test samples via MOS scores. We also present a comparison of how computationally expensive these methods are. Our results show that while speech-driven visual counterpart generation is deemed more natural, the text-driven approach is computationally less demanding and can be used for real-time audiovisual speech synthesis. Also, according to the results all the presented models help to improve the clarity of synthesized speech in noisy conditions.