Internet has undergone a remarkable evolution, reshaped numerous aspects of our world, and left an indelible mark on society. The Internet has unveiled a new era of rapid and accessible interaction, redefining how we connect and converse. In this paper, our objective is to explore various methodologies employed for the conversion of speech into text. The paper embarks on a journey through the historical evolution of audio-to-text conversion, tracing its roots from rudimentary systems to the current state-of-the-art models. This review dissects what are features of audio-to-text conversion and how machine learning models and neural networks influence this process. We have explored various tools and evaluation metrics which are available for the same process. This review further dissects the real-world applications of audio-to-text conversion, domains such as transcription services, voice assistants, and accessibility technologies. It explores the challenges and open research questions on this topic.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comprehensive Survey of Audio-to-Text Conversion

  • Aishwarya Parthasarathi,
  • Almas Banu,
  • Ashwini Joshi

摘要

Internet has undergone a remarkable evolution, reshaped numerous aspects of our world, and left an indelible mark on society. The Internet has unveiled a new era of rapid and accessible interaction, redefining how we connect and converse. In this paper, our objective is to explore various methodologies employed for the conversion of speech into text. The paper embarks on a journey through the historical evolution of audio-to-text conversion, tracing its roots from rudimentary systems to the current state-of-the-art models. This review dissects what are features of audio-to-text conversion and how machine learning models and neural networks influence this process. We have explored various tools and evaluation metrics which are available for the same process. This review further dissects the real-world applications of audio-to-text conversion, domains such as transcription services, voice assistants, and accessibility technologies. It explores the challenges and open research questions on this topic.