Machine translation is getting more and more attention in the research community that deals with natural language processing, language resources and language technologies. It is considered to be one of the most important disruptive technologies with immense implications and benefits for mankind. Closely related is the field of speech technologies that enable tasks, such as automatic speech recognition and speech generation. Both machine translation and automatic speech recognition are explored in this research. The main goal of this paper is to examine the possibilities and obstacles of combining automatic speech recognition with machine translation in a web-based audio-video environment, and in a real-time setting in the sports domain that covers football matches for the purpose of creating a multilingual dataset. The research is performed for two language pairs, English-Arabic and English-Croatian. Captions from videos that contain live sports comments were automatically generated by an automatic speech recognition approach, then machine-translated by a popular online machine translation service, and afterwards edited in three distinct processing phases that considered different aspects of human involvement. Quality evaluations are performed by native speakers with regard to the criterion of usability, and by applying BLEU, the most prominent automatic machine translation quality metric today.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Creating a Multilingual Dataset in Arabic and Croatian from Sports Videos Through a Data Processing Pipeline Combining ASR and MT

  • Wajdi Zaghouani,
  • Sanja Seljan,
  • Ivan Dunđer,
  • Rashid Yahiaoui,
  • Amer Al-Adwan

摘要

Machine translation is getting more and more attention in the research community that deals with natural language processing, language resources and language technologies. It is considered to be one of the most important disruptive technologies with immense implications and benefits for mankind. Closely related is the field of speech technologies that enable tasks, such as automatic speech recognition and speech generation. Both machine translation and automatic speech recognition are explored in this research. The main goal of this paper is to examine the possibilities and obstacles of combining automatic speech recognition with machine translation in a web-based audio-video environment, and in a real-time setting in the sports domain that covers football matches for the purpose of creating a multilingual dataset. The research is performed for two language pairs, English-Arabic and English-Croatian. Captions from videos that contain live sports comments were automatically generated by an automatic speech recognition approach, then machine-translated by a popular online machine translation service, and afterwards edited in three distinct processing phases that considered different aspects of human involvement. Quality evaluations are performed by native speakers with regard to the criterion of usability, and by applying BLEU, the most prominent automatic machine translation quality metric today.