Currently, there are numerous approaches to tackle the problem of recognizing table structures in unstructured documents, in which deep learning stands out as an efficient one. However, applying the deep learning algorithms in real-world scenarios poses significant challenges due to the diversity of tables found in documents evolving daily. Therefore, researching and comparing various TSR (TSR) models for identifying tables in practical documents is not only academically significant but also holds practical implications for applying Artificial Intelligence (AI) to real-world scenarios. Discovering a suitable or superior model for table recognition can help research groups and businesses save time, costs, and optimize product development pathways. In this paper, we perform a comparative study between the two state-of-the-art models through the years, namely Local and Global Pyramid Mask Alignment (LGPMA) and Table Transformer (TATR), using two datasets that we created to represent practical tables. These datasets include the Borderless Table dataset, representing tables commonly found in research articles, and the Vietnamese Tax Liability dataset. To evaluate the performance of the above models, we used the Tree-Edit-Distance-Based Similarity metric, considering only the table structure (TEDS-Struc.). Finally, we draw conclusions and suggest some recommendations for the practical applications of these methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Study of LGPMA and Table Transformer in Table Structure Recognition

  • Tang Van Nguyen,
  • Dang Hai Bui,
  • Phuong Anh Nguyen

摘要

Currently, there are numerous approaches to tackle the problem of recognizing table structures in unstructured documents, in which deep learning stands out as an efficient one. However, applying the deep learning algorithms in real-world scenarios poses significant challenges due to the diversity of tables found in documents evolving daily. Therefore, researching and comparing various TSR (TSR) models for identifying tables in practical documents is not only academically significant but also holds practical implications for applying Artificial Intelligence (AI) to real-world scenarios. Discovering a suitable or superior model for table recognition can help research groups and businesses save time, costs, and optimize product development pathways. In this paper, we perform a comparative study between the two state-of-the-art models through the years, namely Local and Global Pyramid Mask Alignment (LGPMA) and Table Transformer (TATR), using two datasets that we created to represent practical tables. These datasets include the Borderless Table dataset, representing tables commonly found in research articles, and the Vietnamese Tax Liability dataset. To evaluate the performance of the above models, we used the Tree-Edit-Distance-Based Similarity metric, considering only the table structure (TEDS-Struc.). Finally, we draw conclusions and suggest some recommendations for the practical applications of these methods.