<p>Palm leaf manuscripts, celebrated for their historical and cultural significance, pose considerable challenges for digital analysis due to their intricate script structures and age-related degradation. Existing studies frequently highlight the difficulties in assembling realistic datasets and generating accurate ground truth labels, as interpreting ancient scripts requires extensive human expertise and substantial computational resources. To address these challenges, this paper proposes a low-intervention, dual-loop iterative framework designed to efficiently expand datasets and enhance glyph classification in palm leaf manuscript analysis. The framework comprises two primary stages: preprocessing and classification. In the preprocessing stage, state-of-the-art methods are employed for text-line detection, glyph extraction, and synthetic data generation, significantly reducing reliance on manual annotation. The classification stage introduces tailored enhancements to vision transformers (ViTs), incorporating CNN-based feature extraction, Dynamic Stride Shift Patch Tokenization (DS-SPT), and Multi-Scale Locality Self-Attention (MS-LSA). These methods enhance the model’s flexibility and adaptability to the unique characteristics of palm leaf datasets. Moreover, the classification phase facilitates the generation of labels for newly extracted and generated datasets, employing an iterative process to progressively refine model performance. In our experiments, we evaluate the framework using both the ICFHR 2018 palm leaf collection and newly extracted datasets. The experimental results demonstrate improvements in classifying complex glyphs, providing a scalable and efficient solution for low-resource historical document analysis. This framework establishes a foundation for advanced research in the preservation and study of ancient scripts, enabling long-term accessibility and conservation of these cultural heritage documents with minimal human intervention.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Low-Intervention Dual-Loop Iterative Process for Efficient Dataset Expansion and Classification in Palm Leaf Manuscript Analysis

  • Nimol Thuon,
  • Jun Du,
  • Panhapin Theang,
  • Ratana Thuon

摘要

Palm leaf manuscripts, celebrated for their historical and cultural significance, pose considerable challenges for digital analysis due to their intricate script structures and age-related degradation. Existing studies frequently highlight the difficulties in assembling realistic datasets and generating accurate ground truth labels, as interpreting ancient scripts requires extensive human expertise and substantial computational resources. To address these challenges, this paper proposes a low-intervention, dual-loop iterative framework designed to efficiently expand datasets and enhance glyph classification in palm leaf manuscript analysis. The framework comprises two primary stages: preprocessing and classification. In the preprocessing stage, state-of-the-art methods are employed for text-line detection, glyph extraction, and synthetic data generation, significantly reducing reliance on manual annotation. The classification stage introduces tailored enhancements to vision transformers (ViTs), incorporating CNN-based feature extraction, Dynamic Stride Shift Patch Tokenization (DS-SPT), and Multi-Scale Locality Self-Attention (MS-LSA). These methods enhance the model’s flexibility and adaptability to the unique characteristics of palm leaf datasets. Moreover, the classification phase facilitates the generation of labels for newly extracted and generated datasets, employing an iterative process to progressively refine model performance. In our experiments, we evaluate the framework using both the ICFHR 2018 palm leaf collection and newly extracted datasets. The experimental results demonstrate improvements in classifying complex glyphs, providing a scalable and efficient solution for low-resource historical document analysis. This framework establishes a foundation for advanced research in the preservation and study of ancient scripts, enabling long-term accessibility and conservation of these cultural heritage documents with minimal human intervention.