The identification of fraud in health insurance claims is important because of the huge cost implications involved. This paper aims to determine the applicability of random forest classifiers (RFC), support vector machines (SVM), and K-nearest neighbors (KNNs) in detecting fraudulent claims. Data cleaning and data preprocessing were done, and new feature engineering was done from various datasets such as patient dataset, inpatient dataset, and outpatient dataset. Nevertheless, if the first was equipped with a pretend AUC of 0 and an accuracy of 71%, other comparisons indicated that the Random Forest Classifier outperformed the other two models, with achieved accuracies of 82% for SVM and 80% for KNN. These were the diagnosis index, and the Chronic Disease Index. Using threefold cross-validation, the method involving the use of the precision/recall graph, ROC curve, and calibration curve was used to establish the model’s capacity to differentiate between genuine and fake claims. Based on the findings of this study, it is recommended that Random Forest Classifiers are suitable for this task. Subsequent studies should aim at using higher levels of analysis and develop these models while using larger samples. The two objectives of this research are to design advanced, timely fraud detection instruments, and control mechanisms that are instrumental in cutting costs and enhancing healthcare’s effectiveness. Consequently, this research signifies that machine learning classifiers, particularly Random Forest, hold promise in identifying and eradicating healthcare fraud to foster improved and less expensive medical service delivery.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unlocking Healthcare Fraud Detection Using Innovations of Machine Learning Strategies

  • Rahul Dattangire,
  • Divya Biradar,
  • Leelkanth Dewangan,
  • Ashish Joon

摘要

The identification of fraud in health insurance claims is important because of the huge cost implications involved. This paper aims to determine the applicability of random forest classifiers (RFC), support vector machines (SVM), and K-nearest neighbors (KNNs) in detecting fraudulent claims. Data cleaning and data preprocessing were done, and new feature engineering was done from various datasets such as patient dataset, inpatient dataset, and outpatient dataset. Nevertheless, if the first was equipped with a pretend AUC of 0 and an accuracy of 71%, other comparisons indicated that the Random Forest Classifier outperformed the other two models, with achieved accuracies of 82% for SVM and 80% for KNN. These were the diagnosis index, and the Chronic Disease Index. Using threefold cross-validation, the method involving the use of the precision/recall graph, ROC curve, and calibration curve was used to establish the model’s capacity to differentiate between genuine and fake claims. Based on the findings of this study, it is recommended that Random Forest Classifiers are suitable for this task. Subsequent studies should aim at using higher levels of analysis and develop these models while using larger samples. The two objectives of this research are to design advanced, timely fraud detection instruments, and control mechanisms that are instrumental in cutting costs and enhancing healthcare’s effectiveness. Consequently, this research signifies that machine learning classifiers, particularly Random Forest, hold promise in identifying and eradicating healthcare fraud to foster improved and less expensive medical service delivery.