Hearing Voices: Human Sound, Aural Perception, and Early Voice Synthesis
摘要
This chapter begins with a brief overview of Hatsune Miku, Yamaha’s most well-known Vocaloid product. A fan-created song from 2014 is examined as a case study, to explore how we recognize and understand human singing. Next, the chapter offers practical definitions of ‘voice’ as a foundation for understanding singing voice synthesis. It provides an overview of the physiology of vocalization and hearing, covering the basics of phonation, resonance, and articulation, including the distinction between harmonics and formants. Additionally, it describes the auditory system, touching on aural effects such as loudness constancy and perceptual completion. The chapter also delves into several historical efforts to understand human utterance through both physical and perceptual models. The primary physical model of voice discussed is Kempelen’s remarkable eighteenth-century Speaking Mechanism, which mechanically synthesized vocal sounds. Helmholtz’s groundbreaking work on acoustic resonance and auditory perception in the nineteenth century introduces the perceptual approach, with Helmholtz resonators and tuning forks illustrating the basic concepts of analysis and synthesis of complex waveforms from sinusoids. The chapter introduces key concepts such as frequency and pitch, and includes a brief discussion on the distinctions between speech and singing. Throughout, the chapter emphasizes that the production and perception of voice are both inextricably linked and deeply embodied.