Nigerian native languages can still be classified as under-resourced concerning the text and speech resources required for technology development. This limitation strikes all human endeavor facets, including healthcare. The study aims to develop multilingual Nigerian speech resources (English, Yorùbá, and Pidgin languages speech) for antenatal orientation. We collected the dataset from LTH in six (6) different Antenatal clinics using an 8 GB Sony digital voice recorder Midget system, along with a language expert. Each clinic had about an hour of orientation before the nurses began to attend to the pregnant women one after the other. The collected dataset was aligned with the World Health Organisation standard on antenatal. The dataset collected was transcribed into English language and later annotated into Yorùbá (Yorùbá Oyo) and Pidgin languages. The speech version was also developed for the three languages. The project uses the English, Yorùbá, and Pidgin language datasets in text and speech. The word count is 2639, 3202, and 2521 for English, Yorùbá, and Pidgin languages, respectively. These were produced (Speech) in 59880, 70380, and 69840 for English, Yorùbá, and Pidgin languages, respectively. The size of the speech dataset was 15.6, 17.9, and 18.3KB for English, Yorùbá, and Pidgin languages, respectively. The dataset was harvested with a higher level of annotation used as a baseline from under-resourced multilingual Nigeria speech for antenatal orientation and out-of-domain speech data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Developing Nigeria Multilingual Languages Speech Datasets for Antenatal Orientation

  • Sunday Adeola Ajagbe

摘要

Nigerian native languages can still be classified as under-resourced concerning the text and speech resources required for technology development. This limitation strikes all human endeavor facets, including healthcare. The study aims to develop multilingual Nigerian speech resources (English, Yorùbá, and Pidgin languages speech) for antenatal orientation. We collected the dataset from LTH in six (6) different Antenatal clinics using an 8 GB Sony digital voice recorder Midget system, along with a language expert. Each clinic had about an hour of orientation before the nurses began to attend to the pregnant women one after the other. The collected dataset was aligned with the World Health Organisation standard on antenatal. The dataset collected was transcribed into English language and later annotated into Yorùbá (Yorùbá Oyo) and Pidgin languages. The speech version was also developed for the three languages. The project uses the English, Yorùbá, and Pidgin language datasets in text and speech. The word count is 2639, 3202, and 2521 for English, Yorùbá, and Pidgin languages, respectively. These were produced (Speech) in 59880, 70380, and 69840 for English, Yorùbá, and Pidgin languages, respectively. The size of the speech dataset was 15.6, 17.9, and 18.3KB for English, Yorùbá, and Pidgin languages, respectively. The dataset was harvested with a higher level of annotation used as a baseline from under-resourced multilingual Nigeria speech for antenatal orientation and out-of-domain speech data.