Computational Linguistics for Ho Language
摘要
Computational linguistics uses computer science methods to analyze and synthesize language and speech, crucial for understanding human languages. In India, 8.6% of the population is tribal, with different languages, cultures, and customs. English is widely used for communication, but many are deficient in English. Machine translations help overcome language barriers. The Ho language is spoken by the Ho, Kolha, Kol, and Munda tribes. It has a distinct culture, literature, and script. This is the first attempt to create computational linguistics for the Ho language, allowing both Ho and non-Ho speakers to learn and gain knowledge. In our research work, we have prepared a parallel corpus (English-Ho), which contains small sentences and words. The corpus size of the Ho language is twelve thousand five hundred sentence pairs, and four thousand three-hundred word pairs were used for training and testing. We have used Statistical Machine Translation (SMT) and Neural Machine Translation (NMT) to experiment, having an accuracy of around 81%. The accuracy was measured by using the BLEU and ChrF scores. We examined automatic text summarization of Ho language using machine learning algorithms to reduce long texts into small sections of relevant sentences.