Using two languages within one conversation is known as code-switching and poses certain problems for Automatic Speech Recognition (ASR) solutions, especially in multilingual societies. Most basic ASR models sometimes fail to recognize language transitions and have a poor phoneme inventory that leads to high Word Error Rates (WER) and Character Error Rates (CER). Prescribing to low-resource languages such as Hindi and Marathi, this work introduces a novel Hybrid temporal-acoustic Model designed to enhance recognition in code-switched speech. The novel architecture is comprised of multi-scale temporal modelling, adaptive acoustic modelling, and decoding based on reinforcement learning. The system includes a hierarchical encoder, an encoder of a dynamic language embedding generator, a hierarchical attention decoder, and an adversarial language discriminator. These properties, collectively, allow the model to generate linguistic-invariant language embeddings and track the shifts in language smoothly and accurately. In extensive experiments conducted on Hindi-Marathi code-switched speech datasets, we obtained a 45% reduction in WER (down to 18.3 %), a 40% reduction in CER and Language Identification (LID) accuracy higher than 92% while the improvement over the baseline and the state-of-the-art models in averages in 5% for LID. Such observations validate the model’s feasibility for dealing with phonetic variation and sudden shift of languages. This work provides a practical solution for low-resource multilingual ASR and lays down a structure that may be used for further studies on other decentralized language pairs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hybrid Temporal-Acoustic Modeling for Enhanced Code-Switched ASR in Low-Resource Languages: A Study on Hindi and Marathi

  • Hemant Palivela,
  • Meera Narvekar

摘要

Using two languages within one conversation is known as code-switching and poses certain problems for Automatic Speech Recognition (ASR) solutions, especially in multilingual societies. Most basic ASR models sometimes fail to recognize language transitions and have a poor phoneme inventory that leads to high Word Error Rates (WER) and Character Error Rates (CER). Prescribing to low-resource languages such as Hindi and Marathi, this work introduces a novel Hybrid temporal-acoustic Model designed to enhance recognition in code-switched speech. The novel architecture is comprised of multi-scale temporal modelling, adaptive acoustic modelling, and decoding based on reinforcement learning. The system includes a hierarchical encoder, an encoder of a dynamic language embedding generator, a hierarchical attention decoder, and an adversarial language discriminator. These properties, collectively, allow the model to generate linguistic-invariant language embeddings and track the shifts in language smoothly and accurately. In extensive experiments conducted on Hindi-Marathi code-switched speech datasets, we obtained a 45% reduction in WER (down to 18.3 %), a 40% reduction in CER and Language Identification (LID) accuracy higher than 92% while the improvement over the baseline and the state-of-the-art models in averages in 5% for LID. Such observations validate the model’s feasibility for dealing with phonetic variation and sudden shift of languages. This work provides a practical solution for low-resource multilingual ASR and lays down a structure that may be used for further studies on other decentralized language pairs.