Mining and Analyzing Statistical Information from Untranscribed Form Images
摘要
Large amounts of historical structured documents are available in archives around the world. Here, we are interested in automatically extracting statistical information of selected attributes contained in these documents. To this end, we propose a pipeline relying on probabilistic indexing and machine learning to mine and analyze relevant information contained in historical form images. These ideas are assessed on a large collection of images containing almost \(295\,000\) forms with results showing the adequateness of the proposed methods to perform the big-data analytic task considered.