错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Audio-Language Learning

  • Jisheng Bai,
  • Yi Su,
  • Kele Xu

摘要

This chapter explores audio-language learning, a cross-modal paradigm that bridges the gap between acoustic signals and linguistic understanding, mirroring humans’ innate ability to perceive, interpret, and describe sounds through language. This field has experienced unprecedented growth largely propelled by recent advancements in natural language processing and large language models, whose powerful semantic comprehension capabilities have revolutionized multimodal learning approaches. The chapter presents a comprehensive framework that begins with audio-language pre-training methodologies. It examines audio pre-trained models, language pre-trained models, and their multimodal counterparts. Following this, the chapter details downstream transfer learning strategies, including task-specific fine-tuning techniques and the development of more versatile large audio-language models and systems. Additionally, the chapter catalogs crucial audio-language datasets (audio-text paired and audio question-answering data) and benchmarking protocols that collectively establish evaluation standards for this rapidly evolving field. By illuminating the synergistic relationship between audio signal processing and natural language understanding, this chapter provides readers with essential insights into how audio-language learning extends traditional audio analysis capabilities, enabling more intuitive human-machine interaction through natural language interfaces and advancing sophisticated applications from content retrieval and accessibility to audio understanding and generation across diverse domains.