This study emphasizes the importance of natural language processing (NLP) in identifying hate speech on social media when it comes to Vietnamese language by synthesizing many previous papers. A finding from the study is that models like Bidirectional Encoder Representations from Transformers (BERT) and its variations have outperformed Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) in terms of accuracy and reliability. By studying previous research papers on Vietnamese language, it is found that with the complexity and diversity of Vietnamese language such as homonyms, polysemy, slang, regional languages and so on, they are considered as major linguistic challenges that prevent the success in detecting hate speech using NLP tools. However, there is a substantial research gap existing in this field well as in the tools used to process and filter them. To enhance the capacity to identify and reduce harmful speech and support the development of healthier online communities, future research should concentrate on comprehending and resolving these obstacles.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Natural Language Processing (NLP) for Hate Speech Detection in Vietnamese Language: Challenges and Implementation

  • Van Cong Pham,
  • Thair Al-Dala’in

摘要

This study emphasizes the importance of natural language processing (NLP) in identifying hate speech on social media when it comes to Vietnamese language by synthesizing many previous papers. A finding from the study is that models like Bidirectional Encoder Representations from Transformers (BERT) and its variations have outperformed Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) in terms of accuracy and reliability. By studying previous research papers on Vietnamese language, it is found that with the complexity and diversity of Vietnamese language such as homonyms, polysemy, slang, regional languages and so on, they are considered as major linguistic challenges that prevent the success in detecting hate speech using NLP tools. However, there is a substantial research gap existing in this field well as in the tools used to process and filter them. To enhance the capacity to identify and reduce harmful speech and support the development of healthier online communities, future research should concentrate on comprehending and resolving these obstacles.