Automated Enhancement of isiZulu Data Collection for the African Health Research Institute
摘要
IsiZulu, despite being one of the most widely spoken languages of South Africa, is classified as under-resourced. As a consequence, development in Automatic Speech Recognition (ASR) in these languages is slow. Furthermore, data available for training ASR in isiZulu is currently constrained to domains that comprise mainly of content that are government, biblical or news in nature. This poses limitation to developing ASR systems that are more fit-for-purpose. This paper investigates various methods for improving the accuracy of an isiZulu ASR system that is intended to transcribe audio in a domain that does not exist in the pre-trained model. The experimental results in this work shows promising results when applying adaptation techniques using limited domain-specific data. In addition, the paper presents a semi-automated transcription system that includes speech scoring and human-in-the-loop verification of the output transcriptions. The paper also discusses the challenges of adapting ASR models to domain-specific data and presents the contributions and novelty made to the field from the perspective of handling an under-resourced language.