Audio-Related Multimodal Learning
摘要
This chapter delves into the integration of audio with visual and textual modalities, a fundamental component of human perception. Focusing on the intersection of these modalities, we examine their joint application across a wide range of tasks. We begin by categorizing audio into several major sub-domains and introducing the corresponding tasks and applications within each. Next, we summarize relevant datasets and representative models, and systematically review the prevailing modeling paradigms from a structural perspective. Finally, we highlight several key tasks to illustrate recent developments and the challenges that remain in this field.