<p>Detecting fraudulent claims in motor insurance remains a critical challenge due to the severe class imbalance between fraudulent and legitimate cases. This study systematically evaluates a diverse set of resampling strategies—including oversampling, under-sampling, and hybrid methods—in combination with multiple machine learning classifiers. We propose a novel hybrid technique that integrates Center Point SMOTE (CP-SMOTE) with Random Under-Sampling (RUS) to effectively address class imbalance. To evaluate its generalizability, this approach was applied across three publicly available motor insurance datasets, each exhibiting distinct imbalance ratios and feature complexities. Experimental results demonstrate that the CP-SMOTE + RUS combination consistently enhances recall for fraudulent cases while maintaining acceptable levels of precision and the area under the precision-recall curve (AUC-PR) across different model configurations. Notably, on a dataset with a 1:16 fraud-to-non-fraud ratio, the polynomial-kernel SVM achieved a recall of 94.03%. This study validates the robustness of CP-SMOTE + RUS across multiple datasets. The findings underscore the critical impact of class distribution and feature dimensionality on model performance and offer actionable insights for deploying high-recall fraud detection systems in real-world insurance applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing fraud detection in imbalanced motor insurance datasets using CP-SMOTE and Random Under-Sampling

  • Pornpawee Komsrimorakot,
  • Thitirat Siriborvornratanakul

摘要

Detecting fraudulent claims in motor insurance remains a critical challenge due to the severe class imbalance between fraudulent and legitimate cases. This study systematically evaluates a diverse set of resampling strategies—including oversampling, under-sampling, and hybrid methods—in combination with multiple machine learning classifiers. We propose a novel hybrid technique that integrates Center Point SMOTE (CP-SMOTE) with Random Under-Sampling (RUS) to effectively address class imbalance. To evaluate its generalizability, this approach was applied across three publicly available motor insurance datasets, each exhibiting distinct imbalance ratios and feature complexities. Experimental results demonstrate that the CP-SMOTE + RUS combination consistently enhances recall for fraudulent cases while maintaining acceptable levels of precision and the area under the precision-recall curve (AUC-PR) across different model configurations. Notably, on a dataset with a 1:16 fraud-to-non-fraud ratio, the polynomial-kernel SVM achieved a recall of 94.03%. This study validates the robustness of CP-SMOTE + RUS across multiple datasets. The findings underscore the critical impact of class distribution and feature dimensionality on model performance and offer actionable insights for deploying high-recall fraud detection systems in real-world insurance applications.