Faster CNN-Based Layout Analysis of Punjabi Newspapers Using the Custom Dataset
摘要
Layout analysis is an important step in the recognition of text from scanned newspapers. In this paper, we have collected newspapers from the Punjabi Tribune. These newspaper images are then pre-processed and converted into the format for annotation resized to 640 × 640. Then we created the dataset by annotating the images of Punjabi Newspapers. A deep learning-based approach is then used to segment the newspapers into different segments. A list of experiments has been conducted resulting in the accuracy of different classes of newspapers like headlines, text, photographs, advertisements, and title. This method achieves very good accuracy in Punjabi newspapers on different classes.