Historical Document Image Binarization Using Support Vector Machine and Adaptive Threshold Selection
摘要
Segmenting text from severely degraded document images poses a significant challenge due to the considerable variability between the document's background and foreground. Using a single binarization method is inadequate due to the complex degradation factors historical documents experience, including temperature, environment, and paper quality. As a result, an optimal threshold selection process is essential for accurate binarization, which involves separating the foreground text from the background.In this study, we propose a novel approach to binarize historical documents using a combination of support vector machines and a local adaptive threshold selection method. The approach involves initially dividing the input historical document image into multiple subblocks. The features derived from these subblocks are then inputted into an SVM classifier to determine the most appropriate threshold for binarizing each block. Notably, three distinct thresholds are employed for image block binarization.Our method's effectiveness was evaluated using the DIBCO Dataset and compared to various similar techniques. The proposed approach yielded superior results. This effectiveness was demonstrated through testing on two well-known public datasets, DIBCO 2009 and DIBCO 2013, commonly used in document image binarization assessments. The achieved F-Measure for DIBCO 2009 was 89.37%, while for DIBCO 2013, it was 80.31%.