<p>Text data often exhibits class distribution imbalance, causing classifiers to favor the majority class with a larger number of samples, resulting in frequent misclassification of minority class samples. Furthermore, numerical text representations are often highly dimensional. Researchers have proposed several methods employing synthetic minority over-sampling technique with k-nearest neighbors (k-NN) to generate synthetic instances for the minority class by linearly interpolating between a chosen minority instance and its k-NN neighbors. However, k-NN-based approaches may face challenges related to the curse of dimensionality, especially in the case of text oversampling. To address these challenges, we introduce in this paper synthetic genetic oversampling (SYNGO), a novel technique based on genetic algorithms for high-dimensional data that does not rely on neighboring instances. Moreover, SYNGO detects new patterns using the inherent diversity of the majority class by introducing not only the minority instances but also the majority borderline instances into the initial population. We conducted several experiments to verify the performance of the proposed method. The results show SYNGO’s effectiveness in classifying imbalanced text data and that it outperforms other oversampling methods in downstream classification tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Syngo: synthetic genetic oversampling technique for textual data

  • Sihem Nouas,
  • Lamia Oukid,
  • Fatima Boumahdi

摘要

Text data often exhibits class distribution imbalance, causing classifiers to favor the majority class with a larger number of samples, resulting in frequent misclassification of minority class samples. Furthermore, numerical text representations are often highly dimensional. Researchers have proposed several methods employing synthetic minority over-sampling technique with k-nearest neighbors (k-NN) to generate synthetic instances for the minority class by linearly interpolating between a chosen minority instance and its k-NN neighbors. However, k-NN-based approaches may face challenges related to the curse of dimensionality, especially in the case of text oversampling. To address these challenges, we introduce in this paper synthetic genetic oversampling (SYNGO), a novel technique based on genetic algorithms for high-dimensional data that does not rely on neighboring instances. Moreover, SYNGO detects new patterns using the inherent diversity of the majority class by introducing not only the minority instances but also the majority borderline instances into the initial population. We conducted several experiments to verify the performance of the proposed method. The results show SYNGO’s effectiveness in classifying imbalanced text data and that it outperforms other oversampling methods in downstream classification tasks.