错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Zipf Curves and Basic Text Analytics from Untranscribed Manuscript Images

  • Enrique Vidal,
  • Alejandro H. Toselli

摘要

The development of Probabilistic Indexing (PrIx) for large scale historical manuscript collections was originally driven by the need of searching for textual information in large collections of untranscribed text images. The spots that result from the PrIx process are not image transcripts, but they provide very rich probabilistic information about the text rendered in specific regions or locations of the images. This paper presents approaches that exploit this information to go beyond information search applications. In particular we consider text analytics tasks that traditionally require proper textual data such as electronic text. First, for a PrIx-indexed text image collection, word frequency statistical expectations are derived. These expected frequencies are then used to estimate the Zipf curve for the collection. Finally, based on the properties of Zipf curves, we estimate the amount of running words and the size of the lexicon used in the text images considered. Experimental results, reported on several large datasets, show that Zipf curves estimated from text images, accurately match the real curves computed using the ground truth image transcripts. In addition, the parameters derived from these curves (running words and lexicon size) are also fairly accurate approximations to the real values computed from the ground truth plaintext.