Breaking Boundaries: Enhancing Script Identification Using a Learnable MULLER Resizer
摘要
Effective script identification is pivotal in document analysis and recognition, especially when dealing with hybrid multiscript documents at fine-grained levels. Recent advancements in deep learning models have revolutionized document analysis and recognition tasks, leveraging script document images. However, these images often require resizing, which can result in data loss. This paper explores the impact of MULLER, a learnable resizer on a hybrid word-level script identification task using the Multi-lingual and Multi-script Documents In the Wild (MDIW-13) dataset. Our approach integrates MULLER resizer with a pre-module that employs k-means clustering to determine the optimal target size. When jointly trained with MobileNet, it achieves an impressive average accuracy of 98.16%. In summary, our findings underscore the potential of the MULLER resizer to outperform conventional resizers, thereby evaluating script classification performance.