错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sensitive Content Classification

  • Harsha Vardhan Puvvadi,
  • Shyamala L

摘要

In this era of ease of sharing information on the Internet, it has become incredibly easy to share any sort of information online. However, this ease of sharing can come with a great risk of sharing personal or private information, whether knowingly or unknowingly. The potential consequences of compromising information on the Internet can be harmful as it can lead to various forms of online harassment and malpractices. This is why individuals need to be careful about what they share online. A medium is required that can classify the sensitivity of a text to alert the individuals. Many existing approaches classify the text based on the number of sensitive tokens identified. However, this is not enough because these approaches cannot understand the context of the text. In this paper, we proposed a hybrid model leveraging the advantages of CNN, BiLSTM, and multihead attention mechanism, we analyzed the patterns and compared the results provided by standard machine learning and deep learning models, we also discussed the advantages and disadvantages of every model, in extension to do this we also. Our proposed model showed similar to better results than that of the ALBERT model with a significantly much shorter amount of training time.