Table image recognition technology aims to identify and resolve the structure and content of the table from the image. Its process usually includes three steps: table detection, table structure identification and table content identification. Table image content reconstruction is to transform the table image into a structured sequence, in which the structure identification of complex tables is the difficulty of this task. The existing methods are insufficient in identifying the text information, position relationship and logical information of complex cells, and the large computation amount and parameters of the model also become the constraints. In this paper, we propose a new method that introduces the tabular image cell spatial positioning identification branch of the axis attention mechanism, and design the ConvStem backbone network combined with the tabular structure sequence of Transformer to generate branches, and finally embed the OCR module to realize the tabular image content reconstruction. The experimental results show that the proposed method can effectively perform content sequence reconstruction of various kinds of complex tabular images, and conduct comparative experiments on multiple datasets, such as 16% lower than the best model on PubTabNet dataset, while maintaining comparable recognition performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Tabular Image Content Reconstruction Model for Two-Branch Network Design

  • Sun Jun,
  • Liu Jiang,
  • Liu Lijie,
  • Huang Yi,
  • Song Yiting,
  • Zhao Yaxin,
  • Tang Xianxiu,
  • Yang Zhongliao

摘要

Table image recognition technology aims to identify and resolve the structure and content of the table from the image. Its process usually includes three steps: table detection, table structure identification and table content identification. Table image content reconstruction is to transform the table image into a structured sequence, in which the structure identification of complex tables is the difficulty of this task. The existing methods are insufficient in identifying the text information, position relationship and logical information of complex cells, and the large computation amount and parameters of the model also become the constraints. In this paper, we propose a new method that introduces the tabular image cell spatial positioning identification branch of the axis attention mechanism, and design the ConvStem backbone network combined with the tabular structure sequence of Transformer to generate branches, and finally embed the OCR module to realize the tabular image content reconstruction. The experimental results show that the proposed method can effectively perform content sequence reconstruction of various kinds of complex tabular images, and conduct comparative experiments on multiple datasets, such as 16% lower than the best model on PubTabNet dataset, while maintaining comparable recognition performance.