Fraud in the insurance sector presents significant challenges for companies and society, resulting in significant financial losses and impacting pricing strategies. This study addresses the urgent need for effective methodologies to detect fraud, especially when there is data imbalance, with fraudulent cases being less common than legitimate ones. An exhaustive literature review on machine learning models used in fraud detection explicitly focuses on the challenge posed by data imbalance. The study focuses on developing and validating a fraud detection model in the insurance sector, specifically in the realm of automobile insurance claims, addressing the challenge of detecting fraud in unbalanced data environments, where fraudulent cases are less common than legitimate ones. The methodology used to implement the Random Forest Quantile Classifier is detailed, highlighting its ability to optimize sensitivity and specificity in fraud detection by combining the Random Forest technique with quantile classifiers. To evaluate the performance of the proposed model, a case study was conducted using real data provided by an insurance company. The performance of the Random Forest Quantile Classifier was compared with other machine learning methods. The results revealed that the proposed model outperformed others in fraud detection, achieving an optimal balance between true positive and true negative rates. This research contributes to advancing the understanding and refinement of fraud detection techniques, providing valuable insights for professionals and researchers dedicated to financial data analysis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Models for Insurance Fraud Detection: Dealing with Unbalanced Data

  • Patricia Carracedo,
  • David Hervás

摘要

Fraud in the insurance sector presents significant challenges for companies and society, resulting in significant financial losses and impacting pricing strategies. This study addresses the urgent need for effective methodologies to detect fraud, especially when there is data imbalance, with fraudulent cases being less common than legitimate ones. An exhaustive literature review on machine learning models used in fraud detection explicitly focuses on the challenge posed by data imbalance. The study focuses on developing and validating a fraud detection model in the insurance sector, specifically in the realm of automobile insurance claims, addressing the challenge of detecting fraud in unbalanced data environments, where fraudulent cases are less common than legitimate ones. The methodology used to implement the Random Forest Quantile Classifier is detailed, highlighting its ability to optimize sensitivity and specificity in fraud detection by combining the Random Forest technique with quantile classifiers. To evaluate the performance of the proposed model, a case study was conducted using real data provided by an insurance company. The performance of the Random Forest Quantile Classifier was compared with other machine learning methods. The results revealed that the proposed model outperformed others in fraud detection, achieving an optimal balance between true positive and true negative rates. This research contributes to advancing the understanding and refinement of fraud detection techniques, providing valuable insights for professionals and researchers dedicated to financial data analysis.