Large Language Models for Table Content Extraction from Annual Reports
摘要
Our study delves into the utilization of large language models (LLMs) for the extraction of content from tables in annual reports. The study aims to assess the efficacy of LLM in accurately extracting information from annual reports, particularly from tabular data such as balance sheets. The research design involves the analysis of ten annual reports using three different LLMs, with the results recorded in a matrix to evaluate the accuracy of each model. The research paper also evaluates if there are questions that can be answered better than others and if the used LLM has an impact on this. Furthermore, the paper discusses the selection of LLMs for content extraction, data pre-processing, and the evaluation of results. The findings of this research have the potential to inform future developments in LLM and guide the selection of suitable models and methodologies for similar tasks, extending the applicability beyond annual reports to other documents containing tables.