Enhancing Automatic Speech Recognition Systems for Amazigh Language Through Data Augmentation
摘要
Amazigh language faces the challenge of limited training data, a significant obstacle for automatic speech recognition (ASR) researchers. This study addresses this issue by exploring alternative solutions to data scarcity. Experiments were conducted using the Amazigh-Tarifit database, which comprises 18 isolated words in each class, with at least 100 audio repetitions. The dataset, meticulously curated, features recordings from indigenous Moroccan speakers within the Northern Moroccan Rif region. The investigation assesses the impact of data augmentation on ASR, employing three augmentation techniques noise addition, pitch shift, and time stretch on three distinct model architectures (1D CNN, GRU, and 1D CNN LSTM). The results reveal heightened accuracy rates when models are trained on the augmented dataset. Notably, the GRU model achieves a substantial accuracy boost, reaching 97% with augmentation compared to 91% without. These findings affirm the positive influence of data augmentation in enhancing the recognition capabilities of ASR models tailored for the Amazigh language. This work offers valuable insights for overcoming data limitations in underrepresented languages.