Semantic Multi-concept Annotation for Tabular Data in Financial Documents
摘要
Tables in financial documents provide structured data for various analyses, such as the company’s financial health. However, their heterogeneous structures complicate data extraction and narrow the scopes of the analyses. Semantic annotation solves this problem by standardizing the meanings of tabular data, making it fully structured and machine-readable. Although previous research has explored and enhanced semantic annotation, they mainly focus on singular or hierarchical concepts within a table cell, which is insufficient to annotate financial filings. Therefore, we present a more challenging task of annotating multiple non-hierarchical concepts in financial tables. This new task requires a model to identify different concepts describing a table cell. We created a dataset of 10,000 samples and benchmarked seven language models through prompting and fine-tuning. The results demonstrate the challenges of the task, even for large language models, offering the opportunity for future research.