Deep Learning-Based Language Identification in Code-Mixed Text
摘要
Language identification (LID) research is a significant area of study in speech processing. The construction of a language identification system is highly relevant in the Indian context, where almost every state has its language, and each language has many dialects. Social networking is becoming increasingly important in today’s social media platforms for people to convey their opinions and perspectives. As a result, it might be challenging to distinguish between specific languages in a multilingual nation like India. Data was gathered from publicly accessible Facebook postings and tagged with a code-mixed data tag created for this study. This study uses deep learning techniques to identify languages at the word level, where the encoded content may be in Hindi, Assamese, or English. Convolutional neural networks (CNNs) and long short-term memory (LSTM), two deep neural techniques, are compared to feature-based learning for this task. According to the finding, CNN has the best language identification performance with an accuracy of 89.46%.