An Advanced Unimodal Approach to Solving Complex Document Layout Analysis Tasks
摘要
As an important component of document understanding, document layout analysis holds significant development potential. Existing methods predominantly leverage multi-modal features derived from pre-training to enhance performance in layout analysis tasks. However, despite recognizing the multi-modal nature of document images, these methods face challenges in balancing textual and visual features. Additionally, multi-modal approaches often entail high computational overhead and practical deployment challenges. Motivated by these issues, we propose a unimodal method specifically designed for document layout analysis, referred to as DocSalience DETR. When fine-tuned for layout analysis tasks, DocSalience DETR demonstrates superior localizability and recognition capabilities. The proposed method emphasizes features that have undergone channel mapping by pre-processing them in advance, enabling the subsequent encoder to extract more refined features. Furthermore, we select feature scales more suitable for layout analysis to perform token fusion, thereby mitigating query semantic misalignment issues. Experimental results indicate that our method achieves significantly higher recall compared to existing approaches and establishes a new state-of-the-art benchmark on the