错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Tamil News Classification Using LSTM

  • R. Ahila Priyadharshini,
  • P. Kasiviswanathan,
  • B. N. Thana Vignesh

摘要

Massive amounts of news data have been published. Since news is published on numerous websites or tweets (500 million tweets per day), categorizing it is a difficult problem that is not amenable to human categorization. As a result, a novel method is developed to categorize Tamil news items utilizing text preprocessing techniques such stemming, parts of speech tagging, and stopword removal. The proposed method uses Long Short-Term memory (LSTM) to categorize news stories into many categories, such as politics, sports, entertainment, and more. Using text preprocessing techniques on the news items improves the categorization process efficiency. Inflected words are reduced to their word stem using the stemming procedure and stopwords are eliminated using the created stopwords removal procedure. The study’s findings demonstrate that the suggested system works better than conventional approaches and can classify news articles with a high degree of accuracy on Tamilmurasu dataset. Researchers, news organizations, and anyone interested in automatically organizing and analyzing Tamil news stories might find this new method interesting.