<p>The Available Information contained in a hydrological dataset represents the meaningful information per timestep that enters a rainfall-runoff model. To estimate this quantity, Independent Component Analysis (ICA) must be performed to transform the original observations into statistically independent signals. However, ICA algorithms may detect many sets of nearly independent signals suitable for this transformation, posing the question of which set is the optimal one. In the present paper, it is proposed that the components of the optimal set must share the least total pairwise mutual information among all sets since the mutual information of a pair of components is a measure of their statistical independence. This novel approach of estimating the available information of a dataset is applied to five basins in Thessaly, Greece and it is compared to the standard approach of taking the average value occurring from multiple ICA runs. It is illustrated that discarding ICA solutions with high total pairwise mutual information shared between their components stabilizes the estimator of Available Information and increases its precision. This comes with the cost of additional computation time since multiple evaluations of bivariate mutual information are needed.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Quantification of the Available Information of Hydrological Datasets: Stabilizing the Estimator of Multivariate Mutual Information

  • Evangelos Findanis,
  • Athanasios Loukas

摘要

The Available Information contained in a hydrological dataset represents the meaningful information per timestep that enters a rainfall-runoff model. To estimate this quantity, Independent Component Analysis (ICA) must be performed to transform the original observations into statistically independent signals. However, ICA algorithms may detect many sets of nearly independent signals suitable for this transformation, posing the question of which set is the optimal one. In the present paper, it is proposed that the components of the optimal set must share the least total pairwise mutual information among all sets since the mutual information of a pair of components is a measure of their statistical independence. This novel approach of estimating the available information of a dataset is applied to five basins in Thessaly, Greece and it is compared to the standard approach of taking the average value occurring from multiple ICA runs. It is illustrated that discarding ICA solutions with high total pairwise mutual information shared between their components stabilizes the estimator of Available Information and increases its precision. This comes with the cost of additional computation time since multiple evaluations of bivariate mutual information are needed.