From Daisy to Miku: Digital Voices in the Information Age
摘要
This chapter gives an account of how Arthur C. Clarke found inspiration for the character of HAL9000, the infamous computer intelligence, in the world’s first digitally synthesized singing, which he encountered at Bell Labs in the early 1960s. The chapter then gives a brief overview of two key mathematical concepts, the Nyquist limit and the Fourier series, which both facilitated the development of digital audio and advanced voice synthesis methods. Several notable singing voice synthesis systems from the 1980s through the early 2000s are introduced leading to a detailed focus on the development of Vocaloid. The Vocaloid synthesis engine up to Vocaloid 6 utilized concatenative synthesis, with voicebanks (also known as voice libraries, databases, and vocal fonts) provided by several companies. The chapter distinguishes between audio sampling and signal sampling in order to clarify how concatenative synthesis, although relying on large databases of recordings, is different from audio sampling. A case study of the first song released using Vocaloid singing synthesis is examined as a case study to highlight early limitations of the product’s sound. The chapter compares Version 1 and Version 2 voicebanks, designed and marketed by two difference companies. It concludes with an examination of design aspects of the Vocaloid’s first character voice, Hatsune Miku, emphasizing the importance of the product’s nonmusical elements.