<p>Ottoman archives contain an extensive collection of documents; however, their utilization remains limited due to labor-intensive manual keyword searches and the small number of individuals capable of reading Ottoman. In this study, an object detection-based approach is proposed for the detection and recognition of characters in Ottoman handwritten and printed documents. To train and test the Ottoman character recognition network, two datasets have been created, one containing handwritten documents and the other containing printed documents. The created datasets have been organized to construct various scenarios for evaluating the performance of the Ottoman character recognition network in different situations. Character detection and recognition performance of the proposed method for these scenarios has been compared with other object detection methods in the literature, namely Faster R-CNN and SSD methods. To investigate the impact of the components of the proposed method on the character recognition performance, a series of ablation studies were performed and then the findings were analyzed. Our method has shown a higher character detection and recognition performance than the other two methods in various scenarios. Particularly, for the scenario trained with document images from both datasets, the proposed method achieved weighted Average Precision values of 98.61 % and 99.79 % for handwritten and printed test images, respectively, demonstrating a much superior character detection and recognition success compared to the other two methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Object Detection-Based Character Recognition Method for Ottoman Handwritten Documents

  • Ali Alper Demir,
  • Ufuk Özkaya

摘要

Ottoman archives contain an extensive collection of documents; however, their utilization remains limited due to labor-intensive manual keyword searches and the small number of individuals capable of reading Ottoman. In this study, an object detection-based approach is proposed for the detection and recognition of characters in Ottoman handwritten and printed documents. To train and test the Ottoman character recognition network, two datasets have been created, one containing handwritten documents and the other containing printed documents. The created datasets have been organized to construct various scenarios for evaluating the performance of the Ottoman character recognition network in different situations. Character detection and recognition performance of the proposed method for these scenarios has been compared with other object detection methods in the literature, namely Faster R-CNN and SSD methods. To investigate the impact of the components of the proposed method on the character recognition performance, a series of ablation studies were performed and then the findings were analyzed. Our method has shown a higher character detection and recognition performance than the other two methods in various scenarios. Particularly, for the scenario trained with document images from both datasets, the proposed method achieved weighted Average Precision values of 98.61 % and 99.79 % for handwritten and printed test images, respectively, demonstrating a much superior character detection and recognition success compared to the other two methods.