Language identification is a challenging task, particularly when dealing with multilingual and noisy text. This research investigates the optimal fusion of language encoding techniques and machine learning algorithms to achieve accurate language detection. Five widely used encoding methods and several cutting-edge machine learning models are evaluated using two diverse datasets with distinct language distributions. The results highlight the superiority of the FastText encoding method when combined with the Extra Trees Classifier model over the traditional Bag of Words encoding and Multinomial Naive Bayes model in language detection tasks. The study emphasizes the critical role of selecting the appropriate encoding method and machine learning model, as certain combinations yield notably superior outcomes. The insights gained from this research have the potential to significantly enhance the effectiveness and efficiency of language identification systems in relevant domains.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FastText and Extremely Randomized Trees for Language Detection: A Powerful Duo for Multilingual Text Analytics

  • Suvankar Das,
  • Uddalak Mitra

摘要

Language identification is a challenging task, particularly when dealing with multilingual and noisy text. This research investigates the optimal fusion of language encoding techniques and machine learning algorithms to achieve accurate language detection. Five widely used encoding methods and several cutting-edge machine learning models are evaluated using two diverse datasets with distinct language distributions. The results highlight the superiority of the FastText encoding method when combined with the Extra Trees Classifier model over the traditional Bag of Words encoding and Multinomial Naive Bayes model in language detection tasks. The study emphasizes the critical role of selecting the appropriate encoding method and machine learning model, as certain combinations yield notably superior outcomes. The insights gained from this research have the potential to significantly enhance the effectiveness and efficiency of language identification systems in relevant domains.