The ongoing fight against financial crime and fraudulent activities is crucial to safeguarding privacy and security. Addressing this issue is critical for protecting global financial integrity and disrupting illicit attempts. Despite the potential of machine learning, its application in fraud detection encounters unique challenges including limited labels, extremely rare fraud cases, and a vast volume of instances. This is a challenging imbalance scenario where learners must rely solely on an exceptionally small set of labeled instances. Active learning and semi-supervised learning have been proposed to handle such situations. However, their effectiveness and usefulness in these scenarios where different resampling techniques can be applied have not been compared. This paper tackles this gap by conducting a comparative analysis of active and semi-supervised learning approaches for fraud detection. We investigate how the imbalance ratio affects their effectiveness and observe their sensitivity to different resampling techniques. Our results show that both frameworks improve fraud detection, even when the imbalance increases. Highly imbalanced datasets are more affected by different settings for requesting expert labeling than the less imbalanced. In less imbalanced domains, active learning tends to outperform others while in highly imbalanced domains, the performance differences are less pronounced. Applying sampling techniques and moderately adjusting the imbalance ratio improves performance across different subsets of the dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Is Expert-Labeled Data Worth the Cost? Exploring Active and Semi-supervised Learning Across Imbalance Scenarios in Financial Crime Detection

  • Bahar Emami Afshar,
  • Paula Branco,
  • Tolga Kurt,
  • Utku Gorkem Ketenci,
  • Hikmet Mazmanoglu

摘要

The ongoing fight against financial crime and fraudulent activities is crucial to safeguarding privacy and security. Addressing this issue is critical for protecting global financial integrity and disrupting illicit attempts. Despite the potential of machine learning, its application in fraud detection encounters unique challenges including limited labels, extremely rare fraud cases, and a vast volume of instances. This is a challenging imbalance scenario where learners must rely solely on an exceptionally small set of labeled instances. Active learning and semi-supervised learning have been proposed to handle such situations. However, their effectiveness and usefulness in these scenarios where different resampling techniques can be applied have not been compared. This paper tackles this gap by conducting a comparative analysis of active and semi-supervised learning approaches for fraud detection. We investigate how the imbalance ratio affects their effectiveness and observe their sensitivity to different resampling techniques. Our results show that both frameworks improve fraud detection, even when the imbalance increases. Highly imbalanced datasets are more affected by different settings for requesting expert labeling than the less imbalanced. In less imbalanced domains, active learning tends to outperform others while in highly imbalanced domains, the performance differences are less pronounced. Applying sampling techniques and moderately adjusting the imbalance ratio improves performance across different subsets of the dataset.