Proposed Hybrid Model of Focused Crawler Based on Images Containing Tables
摘要
The increasing amount of online data has led to a greater demand for web crawlers that can effectively extract information from web pages. One common challenge is dealing with images that contain tables, as traditional text-based crawlers struggle to process them. To address this issue, we have created a specialized hybrid crawler specifically designed to target images with tables. This advanced crawler utilizes sophisticated image processing techniques for accurate data extraction. By combining content-based image retrieval and machine learning algorithms, our approach enables the crawler to recognize and categorize images based on their visual features. Important testing on financial web pages has demonstrated the remarkable accuracy of our crawler in retrieving relevant images containing tables.