Audio Generation
摘要
This chapter introduces audio generation as an emerging and essential task within audio analysis, focusing on synthesizing new audio content rather than analyzing existing signals. This chapter explores audio generation from multiple perspectives, beginning with critical audio features used in generation processes, and then presenting comprehensive coverage of generative models, including autoregressive approaches, autoencoders, generative adversarial networks, normalizing flows, and diffusion models. Special emphasis is placed on text-to-audio generation methods—a rapidly developing area accelerated by recent advancements in natural language processing. The chapter examines both discrete latent space-based approaches and continuous latent space methods, providing insights into state-of-the-art techniques that enable controllable, high-fidelity audio generation across multiple domains, including speech, music, and environmental sounds. Through this comprehensive overview of audio generation methods, the chapter highlights their crucial role in supporting the rapid evolution of modern internet platforms, short-form video content, and diverse multimedia applications that increasingly rely on sophisticated, customizable audio experiences.