A structured document understanding model based on gate mechanism and cross attention
摘要
Structured document understanding aims to read and analyze both textual and structured information in form documents. Existing models lack sufficient integration of cross-modal information, leaving room for improvement in accuracy. To address this issue, this paper proposes a new network GCAF network based on LiLT. In GCAF, a residual gate module is introduced to ensure that important information from each modality is seamlessly inputted into the fusion module. Additionally, Previous methods often rely on simple concatenation for multiple modalities in form documents, which may not facilitate effective cross-modal fusion. this paper introduces a cross-attention mechanism integrated into the layout information side within LiLT. By doing so, the fusion of layout and textual information is achieved more effectively, leading to superior integration of multimodal information. Experimental evaluations on the FUNSD dataset demonstrate that the proposed GCAF model achieves superior performance with an accuracy of 0.8912.