This research assesses the efficacy of diverse classification algorithms in detecting financial restatements using real-world data spanning 2016 to 2020. Comparative analysis involves Decision Tree, Random Forest, Naive Bayes, Logistic Regression, and Support Vector Machine models, evaluated based on F-measure. Impressively, Naive Bayes and Logistic Regression models achieve exceptional F-measures of 0.929 and 0.973, respectively, underscoring their proficiency in identifying restatements. However, the Decision Tree model’s lower F-measure of 0.455 highlights limitations. Intermediate performance is seen in Random Forest and Support Vector Machine models, recording F-measures of 0.917 and 0.877. This study accentuates the pivotal role of precise model selection for robust predictions. Naive Bayes and Logistic Regression emerge as promising choices for restatement detection, while Decision Tree presents constraints. These findings guide future research for adapting and optimizing models in broader financial analysis and fraud detection contexts. Further investigations are warranted to accommodate practical considerations and dataset limitations across diverse scenarios. In summary, this study contributes vital insights into financial restatement detection, emphasizing the significance of meticulous model choice, with results guiding the development of effective frameworks in the realms of fraud detection and financial analysis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting Financial Restatements of Listed Companies in Vietnam Using Data Mining Techniques

  • Nguyen Thi Kim Oanh,
  • Dong Thanh Duong,
  • Nguyen Van Dinh,
  • Ha Manh Hung,
  • Truong Cong Doan

摘要

This research assesses the efficacy of diverse classification algorithms in detecting financial restatements using real-world data spanning 2016 to 2020. Comparative analysis involves Decision Tree, Random Forest, Naive Bayes, Logistic Regression, and Support Vector Machine models, evaluated based on F-measure. Impressively, Naive Bayes and Logistic Regression models achieve exceptional F-measures of 0.929 and 0.973, respectively, underscoring their proficiency in identifying restatements. However, the Decision Tree model’s lower F-measure of 0.455 highlights limitations. Intermediate performance is seen in Random Forest and Support Vector Machine models, recording F-measures of 0.917 and 0.877. This study accentuates the pivotal role of precise model selection for robust predictions. Naive Bayes and Logistic Regression emerge as promising choices for restatement detection, while Decision Tree presents constraints. These findings guide future research for adapting and optimizing models in broader financial analysis and fraud detection contexts. Further investigations are warranted to accommodate practical considerations and dataset limitations across diverse scenarios. In summary, this study contributes vital insights into financial restatement detection, emphasizing the significance of meticulous model choice, with results guiding the development of effective frameworks in the realms of fraud detection and financial analysis.