错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Automatic POS Tagger System for Code Mixed Indian Social Media Text

  • Nihar Jyoti Basisth,
  • Tushar Sachan,
  • Neha Kumari,
  • Shyambabu Pandey,
  • Partha Pakray

摘要

For a range of Natural Language Processing (NLP) applications, including Sentiment Analysis, Sarcasm Detection, Information Retrieval, Question Answering, and Named Entity Identification, text derived from multiple users’ posts and what they comment on social media constitute significant information (IR). All such applications require part-of-speech (POS) tagging to add tag information to the raw text. Code-mixing, a social media user’s natural desire to submit content in multiple languages, presents a difficulty to POS tagging. In addition, sophisticated and freestyle writing increases the intricacy of the issue. For POS tagging of Code-Mixed Indian social media text, a supervised algorithm using Hidden Markov Model (HMM) with the Viterbi algorithm has been developed to address the problem. The suggested system has been trained and tested using publicly accessible social media text in Indian languages (ILs), particularly Bengali, Telugu, English, and Hindi. On the basis of the F-measure, the accuracy of the system-annotated tags have been assessed.