This study presents the solution to recognizing partially broken Khmer characters in printed text documents using image processing techniques and Machine Learning, specifically Deep Learning. It employs quantitative research along with the experimental design as it explores the relationship between different factors and the desired outcome by analyzing empirical data obtained from experimentation with different deep learning architectures. Otsu’ Thresholding method and a morphological operation called Dilation are employed in this research as image processing techniques before passing the output into the optical character recognition process. The study performs a comparative analysis between two Deep Learning architectures, namely ResNet and VGG as feature extraction, in terms of speed and accuracy. On top of that, BiLSTM is used as a sequence modeling technique to perform contextual analysis coming from the feature extraction output before decoding into a series of words to form a sentence. The models are trained on a huge dataset of synthesized text images along with multiple Khmer fonts, where the images have lost partial visual representation to mimic the partially broken text that are found in historical documents. The two proposed models that are a combination of BiLSTM and CNN architecture (ResNet, VGG) outperformed Tesseract in terms of character error rate by 22% for the former and 18% for the latter.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Optical Character Recognition on Partially Broken Khmer Characters in Printed Documents

  • Menghang Hean,
  • Khem Raksa Peou,
  • Neil Ian Cadungog-Uy,
  • Hongly Va,
  • Sa Math,
  • Tharoeun Thap

摘要

This study presents the solution to recognizing partially broken Khmer characters in printed text documents using image processing techniques and Machine Learning, specifically Deep Learning. It employs quantitative research along with the experimental design as it explores the relationship between different factors and the desired outcome by analyzing empirical data obtained from experimentation with different deep learning architectures. Otsu’ Thresholding method and a morphological operation called Dilation are employed in this research as image processing techniques before passing the output into the optical character recognition process. The study performs a comparative analysis between two Deep Learning architectures, namely ResNet and VGG as feature extraction, in terms of speed and accuracy. On top of that, BiLSTM is used as a sequence modeling technique to perform contextual analysis coming from the feature extraction output before decoding into a series of words to form a sentence. The models are trained on a huge dataset of synthesized text images along with multiple Khmer fonts, where the images have lost partial visual representation to mimic the partially broken text that are found in historical documents. The two proposed models that are a combination of BiLSTM and CNN architecture (ResNet, VGG) outperformed Tesseract in terms of character error rate by 22% for the former and 18% for the latter.