Using a Data-Driven Model to Predict Taxpayers Filing False Returns: A Case of Zambia Revenue Authority
摘要
Tax fraud remains a global issue, with significant economic setbacks for many countries, including Zambia. Traditional methods of tackling this challenge often hinge on labeled datasets, which are scarce due to the slow nature of tax audits and the inherent biases in sample selection. To address this data scarcity and offer a more immediate solution, this study introduces an unsupervised approach utilizing K-means clustering alongside anomaly detection techniques. Using an extensive dataset of VAT declarations and associated refund transactions spanning several years, we demonstrate the potential of this method for efficiently identifying potential tax fraud cases. The critical contribution of this paper lies not only in its innovative approach to a longstanding problem but also in its tailored application for the Zambian context. By bypassing the need for exhaustive labeled data, our methodology offers a promising direction for enhancing tax fraud detection capabilities, ensuring a more resilient fiscal landscape for Zambia.