Code-mixing is a common phenomenon in multilingual communities, where speakers use more than one language within a single conversation. A transliteration framework is necessary to accommodate individuals who feel more comfortable primarily using their native tongues for communication in the realm of social media. The proposed framework leverages existing transliteration techniques and incorporates novel approaches to handle the challenges specific to English-Tamil code-mixed data. The word-level language classification was implemented using supervised machine learning models and HMM, with the latter performing better with an accuracy of 75%. Tamil-tagged words from the classifier are passed to the Seq2seq encoder-decoder model for transliteration, and the model achieved an accuracy of 93%. The effectiveness of the framework is evaluated using a corpus of code-mixed data, and the results are compared with existing transliteration methods, demonstrating its efficacy in transliterating code-mixed data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transliteration Framework for Tamil-English Code-Mixed Text

  • K. A. Viswabharathi,
  • K. R. Rithick,
  • T. Thayumanan,
  • K. Nalinadevi

摘要

Code-mixing is a common phenomenon in multilingual communities, where speakers use more than one language within a single conversation. A transliteration framework is necessary to accommodate individuals who feel more comfortable primarily using their native tongues for communication in the realm of social media. The proposed framework leverages existing transliteration techniques and incorporates novel approaches to handle the challenges specific to English-Tamil code-mixed data. The word-level language classification was implemented using supervised machine learning models and HMM, with the latter performing better with an accuracy of 75%. Tamil-tagged words from the classifier are passed to the Seq2seq encoder-decoder model for transliteration, and the model achieved an accuracy of 93%. The effectiveness of the framework is evaluated using a corpus of code-mixed data, and the results are compared with existing transliteration methods, demonstrating its efficacy in transliterating code-mixed data.