<p>We adapt Manifold Mixup theory for accent-robust end-to-end (E2E) Automatic speech recognition (ASR). Accent-variation between a source and target constitutes a domain-mismatch scenario. Manifold Mixup allows cross-domain robustness where a model trained on a source accent generalizes to target accents. We propose a 2-stage training mechanism with manifold mixup using one source accent. Stage 1 is a mixup-enabled cross-entropy based framewise character recognition model. Stage 2 is a Connectionist Temporal Classification (CTC)-loss based E2E ASR model using Stage 1 weights. We show that this model generalizes to unseen accents without any fine-tuning. This is studied for accented English from Indic-TIMIT corpus (6 Indic accents) and Common Voice corpus accent groups UKI (England, Ireland), Oriental (India, Malaysia), NorthAM (USA, Canada), African and ANZ (Australia, New Zealand). This is also studied with another Indian English corpus Svarah, the American English TIMIT corpus and the open audiobook English corpus of Librispeech. The proposed framework, using a Hindi-mixup model, offers absolute gains of around 2% over a non-mixup baseline on unseen test accents.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Accent-robust speech recognition for English in low-resource settings using Manifold Mixup

  • Tirthankar Banerjee,
  • V. Ramasubramanian

摘要

We adapt Manifold Mixup theory for accent-robust end-to-end (E2E) Automatic speech recognition (ASR). Accent-variation between a source and target constitutes a domain-mismatch scenario. Manifold Mixup allows cross-domain robustness where a model trained on a source accent generalizes to target accents. We propose a 2-stage training mechanism with manifold mixup using one source accent. Stage 1 is a mixup-enabled cross-entropy based framewise character recognition model. Stage 2 is a Connectionist Temporal Classification (CTC)-loss based E2E ASR model using Stage 1 weights. We show that this model generalizes to unseen accents without any fine-tuning. This is studied for accented English from Indic-TIMIT corpus (6 Indic accents) and Common Voice corpus accent groups UKI (England, Ireland), Oriental (India, Malaysia), NorthAM (USA, Canada), African and ANZ (Australia, New Zealand). This is also studied with another Indian English corpus Svarah, the American English TIMIT corpus and the open audiobook English corpus of Librispeech. The proposed framework, using a Hindi-mixup model, offers absolute gains of around 2% over a non-mixup baseline on unseen test accents.