错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Empirical Validation of Probabilistic Indexing Methods

  • Alejandro Héctor Toselli,
  • Joan Puigcerver,
  • Enrique Vidal

摘要

The proposed probabilistic framework and most of the specific approaches, algorithms, assumptions, and claims discussed throughout the previous chapters, require empirical validation. This is the main purpose of this chapter. In particular, the most relevant questions that we aim to answer trough the experiments are: 1. As compared with the HTR-oriented formulation proposed in Sec. 3.4, how do the various lexicon-based image-processing-oriented posteriorgram methods discussed in Secs. 3.1 and 3.3.2 perform? This will be studied in Sec. 6.2. 2. As discussed in Chapter 3, text-lines are particularly interesting image regions for indexing purposes. So, the question is, how the different RPs defined in that chapter can advantageously be used under a line-level PrIx paradigm? This will be developed in Sec. 6.3. 3. What is the impact of a language model on PrIx performance? Which general approach is preferable, lexicon-based, or lexicon-free? These questions are tackled in Sec. 6.4. 4. How does the amount of training examples affect PrIx performance? This is studied in Sec. 6.5. 5. Given that both PrIx andHTRuse the same underlying probability distributions, is there a clear correlation between the performance on HTR and PrIx tasks? This topic is examined in Sec. 6.6. 6. Since search is one of the main applications of both PrIx and KWS, how does our PrIx methods compare with state-of-the-art KWS approaches? This is studied at line-level in Sec. 6.7. 7. How much PrIx performance improvement can be expected by using the newer, neural-network-based CRNN optical models with respect to adopting more traditional statistical HMM models? We analyze this question in Sec. 6.8. 8. Can line-oriented PrIx RPs be used to tackle KWS under a segmentation-free paradigm? This is empirically assessed in 6.10.1. 9. Is the approach proposed in Sec. 3.6 to compute the RP for a query image (rather than a textual query) adequate to perform traditional QbE KWS? How does this approach fares with respect to other segmentation-free QbE KWS methods? This is studied in Sec. 6.10.2.