Nowadays, the extraction of actual data from financial reports is a largely manual process, requiring inordinate levels of time and effort. This error-prone procedure requires carefulness and attention to detail, which may lead to errors including entering incorrect data or missing important information if neglected. That being said artificial intelligence research has made some big strides over the last few years. Therefore, developing an automated system to extract data from financial reports is essential. This study aims to build an automated system to address the issues of time, effort, and errors associated with manual extraction methods, applying machine learning techniques and optical character recognition (OCR). This research uses the YOLOV8 model to locate the table part and Surya OCR to fetch financial report image data. Preprocessing methods are also applied to improve the model’s accuracy. The outputs have demonstrated a mean character error rate (MCER) of 0.0223 and a mean word error rate (MWER) of 0.44, exhibiting the system’s high accuracy regarding information extraction. This further confirms the increased effectiveness of an automated method compared to existing methods, underlining its ability to automate the financial report data extraction process and thus assist strategic decisions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Information Extraction from Financial Reports

  • Khong Van Vinh,
  • Thi Duyen Nguyen,
  • Thi Bich Thuy Nguyen,
  • Viet Nguyen Dinh,
  • Doan Truong Cong

摘要

Nowadays, the extraction of actual data from financial reports is a largely manual process, requiring inordinate levels of time and effort. This error-prone procedure requires carefulness and attention to detail, which may lead to errors including entering incorrect data or missing important information if neglected. That being said artificial intelligence research has made some big strides over the last few years. Therefore, developing an automated system to extract data from financial reports is essential. This study aims to build an automated system to address the issues of time, effort, and errors associated with manual extraction methods, applying machine learning techniques and optical character recognition (OCR). This research uses the YOLOV8 model to locate the table part and Surya OCR to fetch financial report image data. Preprocessing methods are also applied to improve the model’s accuracy. The outputs have demonstrated a mean character error rate (MCER) of 0.0223 and a mean word error rate (MWER) of 0.44, exhibiting the system’s high accuracy regarding information extraction. This further confirms the increased effectiveness of an automated method compared to existing methods, underlining its ability to automate the financial report data extraction process and thus assist strategic decisions.