Historical documents are vital for preserving cultural information. To access their content, OCR and HTR technologies transcribe them into text. Challenges arise with handwritten texts due to varied writing styles. This study develops two corpora to train models for 19th-century Greek texts. The first dataset includes printed texts from the “Ellinomnimon” archive, and the second comprises handwritten documents from the Lasithi Demogerontia archives. Moreover, using the Transkribus platform, two models were iteratively refined, enhancing their ability to transcribe Greek historical documents from 1800 to 1870.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Developing Datasets for Training OCR/HTR Models for the Late 19th Century Greek Texts

  • Georgios Roukas,
  • Michalis Sfakakis

摘要

Historical documents are vital for preserving cultural information. To access their content, OCR and HTR technologies transcribe them into text. Challenges arise with handwritten texts due to varied writing styles. This study develops two corpora to train models for 19th-century Greek texts. The first dataset includes printed texts from the “Ellinomnimon” archive, and the second comprises handwritten documents from the Lasithi Demogerontia archives. Moreover, using the Transkribus platform, two models were iteratively refined, enhancing their ability to transcribe Greek historical documents from 1800 to 1870.