A review of distinct machine learning classifiers for healthcare fraud detection
摘要
Healthcare insurance fraud is a major problem, with an estimated $300 billion lost annually in the United States. Machine learning has been explored as a tool for fraud detection for over a decade, but challenges remain, including class imbalance, dataset diversity, and model interpretability. This survey reviews 22 supervised, unsupervised, and semi-supervised classification techniques published between December 2017 and October 2024 which are novel within the healthcare fraud detection domain. Supervised techniques, which make up a majority of the identified works, are divided into deep learning, graph-based, and meta-learning methods. We also find a small set of novel unsupervised and semi-supervised classifiers, which we examine with a focus on model explainability. We find that deep learning is a popular and effective approach for supervised learning, while research gaps remain in areas such as transfer learning and incremental learning for all three classification approaches. Most concerningly, we find few to no works which establish benchmarks across diverse techniques. We assert that addressing these gaps could lead to more effective, transparent, and adaptable fraud detection systems in healthcare insurance.