Proposed Model for Automatic Dialect Classification of Binjhal Language
摘要
India is a multilingual country, where more than 8.6% of tribal people reside in India. The Web’s adoption of text-based Binjhal dialects has presented both fresh opportunities and difficulties for machine learning and Binjhal language processing. Applications like sentiment analysis and machine translation are greatly impacted by the detection of this kind of unstructured data. The nature of dialectal textual data, however, precludes the use of the typical NLP tools created for traditional data. Deep learning methods have shown to be particularly successful at processing dialectal text from social media. Several deep learning models are taken into consideration in this work for the automatic categorization of Binjhal dialectal literature. A reliable and powerful language detection tool is necessary for the development of any NLP tools or technology. Although many high-resource languages have access to these language recognition methods and technologies, low-resource languages do not have as many training datasets, making it more challenging to locate them. Our suggested dataset will enrich Odia, Binjhal, and Sambalpuri and assist academics in creating technologies and tools for this low-resource language.