<p>Historical documents offer invaluable insights into the past, shaping our understanding of the world’s rich tapestry of stories. This paper presents methodologies designed to aid in the transcription of French Belfort birth registers from 1807 to 1919. The approach emphasizes preprocessing techniques, including binarization, skew correction, and text line segmentation, to effectively tackle the challenges arising from varied text styles, marginal annotations, and the combination of printed and handwritten text. We present this archive as a newly developed database, employing a structured methodology with XML tags to ensure accurate formatting and alignment of transcriptions with image elements at both the paragraph and text line levels. Our preprocessing phase demonstrates an accuracy rate of 95.9%, highlighting the effectiveness of our techniques in preserving and facilitating the study of this rich cultural heritage. This work contributes substantially to the field of handwritten text recognition and sets the stage for further advancements in the automated transcription of historical documents. This paper is an extended version of our previous work presented at IMPROVE2024, including additional experimental results and in-depth analysis.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimized Data Structuring and Preprocessing Techniques for Belfort Civil Registers of Birth Transcription

  • Wissam AlKendi,
  • Franck Gechter,
  • Laurent Heyberger,
  • Christophe Guyeux

摘要

Historical documents offer invaluable insights into the past, shaping our understanding of the world’s rich tapestry of stories. This paper presents methodologies designed to aid in the transcription of French Belfort birth registers from 1807 to 1919. The approach emphasizes preprocessing techniques, including binarization, skew correction, and text line segmentation, to effectively tackle the challenges arising from varied text styles, marginal annotations, and the combination of printed and handwritten text. We present this archive as a newly developed database, employing a structured methodology with XML tags to ensure accurate formatting and alignment of transcriptions with image elements at both the paragraph and text line levels. Our preprocessing phase demonstrates an accuracy rate of 95.9%, highlighting the effectiveness of our techniques in preserving and facilitating the study of this rich cultural heritage. This work contributes substantially to the field of handwritten text recognition and sets the stage for further advancements in the automated transcription of historical documents. This paper is an extended version of our previous work presented at IMPROVE2024, including additional experimental results and in-depth analysis.