A review on handwritten text segmentation in Indian languages
摘要
Word segmentation is a crucial phase for handwritten text recognition. When the input image contains several words written into multiple text lines, it is essential to separate or segment the images of individual words. Although several techniques have been proposed, handwritten text segmentation is still a popular research problem due to significant variations in handwriting styles and script-specific challenges. This paper systematically reviews the literature on handwritten text segmentation in Indian languages. The primary objective of the survey is to accumulate the techniques used for the segmentation of handwritten Indian documents. There, we found that a number of techniques, ranging from traditional image processing techniques to modern deep-learning models, have been used for segmentation from Indian handwritten documents. As there are a number of techniques, a comparative analysis of those techniques is necessary to understand their relative efficiency. Therefore, we implemented those techniques and tested their performance using a common dataset. For the comparison, we have taken handwritten images of three different languages Bengali, Hindi, and Tamil. In our experiments, we found that deep learning-based models often performed better than the non-learning-based techniques.