The Covid-19 pandemic has caused a surge in healthcare claims fraud due to increased digitization in the health insurance domain. The dependence on digital systems during the pandemic has led to greater accumulation of data which has in turn spurred newer and innovative frauds. Insurance frauds are a social menace which result in increasing healthcare costs for all, compromising quality of care, breaking the trust of the policyholders with the Insurer and creating legal and regulatory issues for the Insurer. This research aims to identify phenotypes or observable claim characteristics that indicate fraudulent behavior using unsupervised anomaly detection (AD) methods. The experiments are done on the CMS Medicare datasets. They identify unknown phenotypic patterns using the unsupervised AD models, one-class support vector machine (OC-SVM), Isolation Forest (IForest) and local outlier factor (LOF) to identify behavior patterns leading to fraud. These models are ensembled and the resulting outcome is analyzed for feature importance to derive leading characteristics for fraudulent behavior.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Phenotyping Health Insurance Claims Fraud with Unsupervised Anomaly Detection Methods

  • Supriya Seshagiri,
  • K. V. Prema

摘要

The Covid-19 pandemic has caused a surge in healthcare claims fraud due to increased digitization in the health insurance domain. The dependence on digital systems during the pandemic has led to greater accumulation of data which has in turn spurred newer and innovative frauds. Insurance frauds are a social menace which result in increasing healthcare costs for all, compromising quality of care, breaking the trust of the policyholders with the Insurer and creating legal and regulatory issues for the Insurer. This research aims to identify phenotypes or observable claim characteristics that indicate fraudulent behavior using unsupervised anomaly detection (AD) methods. The experiments are done on the CMS Medicare datasets. They identify unknown phenotypic patterns using the unsupervised AD models, one-class support vector machine (OC-SVM), Isolation Forest (IForest) and local outlier factor (LOF) to identify behavior patterns leading to fraud. These models are ensembled and the resulting outcome is analyzed for feature importance to derive leading characteristics for fraudulent behavior.