Topic Modelling and Interpretable Cost Estimation for Medical Insurance Fraud Detection
摘要
Medical insurance incurs significant costs and can be susceptible to fraud or waste. Machine learning approaches to automating fraud detection are becoming commonplace. Real-world pipelines including decision support systems for compliance activities on medical insurance claims may include requirements such human-interpretability and estimates of recoverable costs, to assist with prioritisation and investigation. We previously developed a framework for learning claim contexts and provider roles, incorporating domain knowledge through the insurance item ontology, and ranking providers for audit based on cost differences between similar providers. We extend this by comparing an interpretable pattern-identification cost estimator to the original scoring method and evaluating on a large real-world claims dataset. Compared to our previous cost estimator performance was similar, but with the advantage of immediately interpretable results. Results show incorporating context discovery and domain knowledge into fraud detection algorithms assists identification of comparable providers and generation of interpretable results for subject-matter experts in the decision-support process.