This chapter opens with an account of Vocaloid’s 2019 premier of the AI Hibari voicebank, which was trained on the archive of a renowned twentieth-century Japanese singer. The chapter then offers an overview of artificial intelligence followed by an exploration of the current impact of machine learning on singing voice synthesis (SVS). A discussion of the history and scope of the field of AI covers key technical distinctions, such as predictive versus generative AI and supervised versus unsupervised algorithms. The chapter also introduces deep neural networks (DNN), natural language processing (NLP), and large language models (LLM), while touching on the roles and significance of foundation models and transformer architectures. The chapter compares how different singing voice synthesis systems approach AI training. Several DNN-based SVS systems are presented, including VOCALOID:AI, demonstrated in 2019 with AI Hibari and released for sale in 2022 as Vocaloid 6 with Vocalo Changer, a vocal timbretimbre transfer tool. Concerns surrounding voice cloning are examined through case studies from East Asia between 2018 and 2022. The issue of deepfakes in the USA and Europe is then examined, focusing on the study of Holly Herndon’s Holly+ and Spawning in comparison with Grimes’ Elf.tech. A brief survey of current singing voice synthesis products is undertaken. The chapter concludes with consideration of how to approach ethical concerns about artificial intelligence in singing synthesis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Future Voices to Come: AI Singing After Miku

  • Gretchen Jude

摘要

This chapter opens with an account of Vocaloid’s 2019 premier of the AI Hibari voicebank, which was trained on the archive of a renowned twentieth-century Japanese singer. The chapter then offers an overview of artificial intelligence followed by an exploration of the current impact of machine learning on singing voice synthesis (SVS). A discussion of the history and scope of the field of AI covers key technical distinctions, such as predictive versus generative AI and supervised versus unsupervised algorithms. The chapter also introduces deep neural networks (DNN), natural language processing (NLP), and large language models (LLM), while touching on the roles and significance of foundation models and transformer architectures. The chapter compares how different singing voice synthesis systems approach AI training. Several DNN-based SVS systems are presented, including VOCALOID:AI, demonstrated in 2019 with AI Hibari and released for sale in 2022 as Vocaloid 6 with Vocalo Changer, a vocal timbretimbre transfer tool. Concerns surrounding voice cloning are examined through case studies from East Asia between 2018 and 2022. The issue of deepfakes in the USA and Europe is then examined, focusing on the study of Holly Herndon’s Holly+ and Spawning in comparison with Grimes’ Elf.tech. A brief survey of current singing voice synthesis products is undertaken. The chapter concludes with consideration of how to approach ethical concerns about artificial intelligence in singing synthesis.