YOLOv10-CBRC: A high-precision document image layout analysis model
摘要
Tibetan document layout analysis (DLA) is a crucial aspect of the digitization of Tibetan documents. However, two significant issues persist: (1) the inadequacy of Tibetan document datasets;(2) insufficient utilization of layout information by existing models. To address these challenges, we have constructed a novel dataset, the Aba Tibetan Newspaper Dataset (AbaTND), consisting of 571 color images of Tibetan newspapers, effectively filling a data gap in this field. Additionally, we propose an advanced model, YOLOv10-CBRC (YOLOv10-CBAM-Re-Upsample-Re-SCDown-CRCV3), along with its variant YOLOv10-RC (YOLOv10-Re-Upsample-Re-SCDown-CRCV3), aimed at enhancing information utilization. Building upon YOLOv10 as the baseline model, the proposed model implements three key improvements: (1) the replacement of the Partial Self-Attention (PSA) module with the Convolutional Block Attention Module (CBAM), which enhances perception of channel-wise information and spatial localization of objects,(2) the adoption of dual-branch Re-upsample and Re-SCDown modules (2Re), which facilitates more effective utilization of multi-scale information,(3) the design of a novel classification feature processor, CIBwithResidualCV3 (CRCV3), which improves performance in classification tasks.Experimental results demonstrate that YOLOv10-CBRC achieves a mAP