错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AIE-KB: Information Extraction Technology with Knowledge Base for Chinese Archival Scenario

  • Shiqing Bai,
  • Yi Qin,
  • Peisen Wang

摘要

The current visual information extraction (VIE) methods usually focus only on textual features or image features of the document modality, ignoring the segment-level structural information in the document, thus limiting the accuracy of information extraction. To address the problem of information extraction in the scenario of scanned personnel archival images, the paper proposes an Archive Information Extraction model with Knowledge Base (AIE-KB), which injects knowledge into the model in the coarse relationship prediction module (CRP). Our model undergoes several modules, including entity recognition, CRP, and entity linking, after encoding the features for multi-modal fusion. As detailed in Fig. 2. Specifically, we (1) add new intermediate entity types and (2) design a new relationship prediction head to explicitly exploit structural information in scanned images, solving the problem of nested entity relationships specific to personnel archives. To address the lack of a certain amount of annotated data on personnel files, we also introduce a benchmark dataset of personnel archives in which the “answer” types in each scanned archival image are handwritten. The experimental results show that our model AIE-KB achieves excellent results in two tasks with different datasets, demonstrating the effectiveness and robustness of the model. All metrics for both tasks exceeded the baseline model, with the F1 metric for entity recognition on the Visually rich scanned Images in the personnel Archive (VIA) (from 84.36 to 94.61) and entity linking on VIA (from 69.79 to 87.94).