错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data augmentation and adversary attack on limit resources text classification

  • Fernando Sánchez-Vega,
  • A. Pastor López-Monroy,
  • Antonio Balderas-Paredes,
  • Luis Pellegrin,
  • Alejandro Rosales-Pérez

摘要

Data Augmentation and Adversary Attack in text are complex techniques based on the generation of new instances. This is performed by introducing some variations such as lexical and syntactic changes in previously known text examples. In both techniques, it is mandatory to preserve the general semantic meaning of the text in order to boost or to mislead the classifier in each case. The instance generation on data augmentation is especially important in a lack of data scenario such as limited-resource languages. Using the new instances as training samples could overcome data scarcity problems. In this paper, we adapt four textual instance generation methods used in some Data Augmentation and Adversary Attack methods, and we propose two more methods in order to be used in limited-resource languages environments. We have empirically quantified how much damage the Adversary Attack techniques can cause to the textual classification’s performance in low-resource scenarios. Furthermore, we explore the use of data augmentation with adversarial attack strategies to increase the robustness of classification models against the adversary attacks.