Auto-Insurance Fraud Detection Using Machine Learning Classification Models
摘要
This work explored six machine learning algorithms: Extreme Gradient Boosting (XGBoost), Logistic Regression, Random Forest, Decision tree, Support Vector Machine (SVM), and Naïve Bayes to determine the best algorithm for detecting insurance fraud. The following were used to evaluate the six models: Confusion matrix, Accuracy, Precision, Recall, and F1-measure. The result showed that Random Forest outperformed the others in terms of accuracy. Extreme Gradient Boosting (Xgboost) had the highest precision and F1-measure scores, while the Decision Tree had the highest Recall score. Although two methods (Analysis of Variance (ANOVA) and Random Forest Classifier) were compared to determine the best feature selection, the significant features were selected using the Random Forest classifier because of the many benefits of using this method. The results of this study will be beneficial to insurance companies, stakeholders and policyholders.