错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sentence Level Language Classification of Malayalam-English Mixed-Language Text

  • C. P. Afsal,
  • K. S. Kuppusamy

摘要

Language classification plays a vital role in multilingual text analysis, especially in code-mixed scenarios where multiple languages are combined within a single sentence. User-generated content in social media is loaded with such content. In this work, we present a comprehensive study on sentence-level language classification of Malayalam-English mixed-language text using BERT (Bidirectional Encoder Representations from Transformers) and its variant models. We experimented with various BERT-based models to explore their effectiveness in accurately identifying the language in mixed-language sentences. After conducting experiments on our custom dataset, we found that the BERT and DistilBERT language models, out of all the variants we tested, achieved the highest accuracy of 99.76%. The findings of this study shed light on the effectiveness of BERT-based models for mixed-language text classification tasks.