Development of Parallel Speech Data Repository for Ho Language
摘要
The Ho language, which belongs to the Austroasiatic language of Munda family, is their primary means of communication with the Ho tribe. The Ho tribe is an Indigenous community that primarily inhabits the Indian states of Odisha, Jharkhand, West Bengal, Assam, and Chhattisgarh. Their population is 5 million. Warang Chiti is the script for writing Ho language. Pandit Lako Bodra discovered Warang Chiti Script in 1954. Speech recognition systems are a necessity in our modern society. They provide accessibility for individuals with disabilities, enhance productivity through hands-free operation, enable multilingual communication, support voice control and automation, improve user experiences, offer valuable data insights, enhance driving safety, and advance natural language processing capabilities. Developing an automatic speech recognition (ASR) system for the Ho community involves collecting and pre-processing speech data, followed by training a specialized model for accurate speech recognition and transcription in the Ho language. After collecting the speech data from Ho community and pre-processing the existing data, will prepare an ASR for Ho community.