Machine Learning Analysis on Hate Speech Against Asians
摘要
Racism is a social issue that is being more and more addressed over the past years. However, prejudice against Asians is still present and increasing, especially in social networks. This could be more explicitly observed during the Covid-19 pandemic and the start of Tokyo 2021 Olympics. In order to draw attention to Asian racism, this study explores Natural Language Processing and Machine Learning techniques to identify comments with harmful messages on social media such as X (formerly Twitter). For this, sentiment and elements of micro-aggression on tweets were analyzed in an effort for the trained models to grasp the context and combination of words that indicates the presence or not of racial slurs. Among the trained models, including Bayes, LSTM and CNN, the latter presented the better results and was later used for the development of a Twitter bot, able to consult whether or not any given thread had a tendency to racism. Thus, by the end of this study, racism messages classification was proven to be possible, opening possibilities to deepening on this subject.