LD-DOC: Light-Weight Domain-Adaptive Document Layout Analysis
摘要
We propose the LD-DOC, a lightweight Document Layout Analysis (DLA) model specifically designed to address the challenge of accurately partitioning document regions under limited data conditions. The LD-DOC model effectively utilizes information from various scale visual features, enhancing its adaptability to feature distributions in scenarios with limited data and thereby improving the accuracy of document region partitioning. Specifically, our model incorporates a feature fusion module comprising a Shallow Feature Enhancement Path (SFEP) and a Cross-Fusion Path (CFP). The SFEP employs a 2D-Discrete Wavelet Transform (2D-DWT) to capture edge features at different scales, which enhances the model’s ability to perceive subtle variations and structural information in visual features. This enhancement is crucial for adapting to the nuanced requirements of limited data environments. On the other hand, the CFP uses a Local-Fusion Attention mechanism(LFA) to capture Discrepancy information adaptively among different scales. This approach reduces the model’s sensitivity to scale variations and significantly improves its generalization capabilities across diverse document layouts. Furthermore, we introduce the ISCAS-CLAD , a specialized small-scale Chinese Document Layout Analysis Dataset, to demonstrate the effectiveness of our model. Through rigorous testing on ISCAS-CLAD and the PubLayNet datasets, LD-DOC has shown a notable improvement in mean Average Precision (mAP) accuracy, outperforming baseline models by 2.2 \(\%\) and 1.5 \(\%\) , respectively. These results highlight LD-DOC’s state-of-the-art performance, particularly in challenging data-limited environments, and underscore its potential for practical applications in DLA.