An Improved Skew Detection and Correction Method for Bangla Handwritten Document Using Orthogonal Regression and Connected Component Analysis
摘要
Handwritten document images need to be processed carefully to understand meaningful information. For useful collection and future reference, nowadays these documents are now protected in digital library. Historical records, cultural artifacts, and literary assets in Bangla scripts are precious things which should be stored in digitized format so that it can’t be lost. However, character recognition from document images is challenging due to skewness and sometimes it is unmanageable to retrieve information in text format. Moreover, skew rate can deteriorate according to the handwriting style of non-printed document image, especially for Bangla handwritten document image. For this document, characters are not easily separable, but words are discreet with each other. This research addressed this issue with the application of Connected Component Analysis. Although the skew rate can be measured from connected component analysis and bounding box approaches, orthogonal regression is applied to acquire the best fitting line. Experimental results show that the proposed method provides 98.52% accuracy which indicates the effectiveness of the proposed scheme.