Analyzing Inferential Reproducibility
摘要
Reproducibility of experimental results is one of the fundamental pillars of scientific research. If neither a reliable nor significant evaluation result can be obtained when replicating an experiment, the whole methodological foundation of the research result becomes questionable, even casting doubt on its validity. In this chapter, we will show how to apply the reliability and significance tests introduced in the previous chapters to analyze the reproducibility of research results. The flexibility of the model-based tests allows us to give this analysis an important twist: Instead of following the current tendency to remove measurement noise in order to enforce reproducibility under the exact same training conditions, we aim at a notion of inferential reproducibility (Goodman et al., 2016). This concept embraces certain types of nondeterminism as inherent and irreducible conditions of measurement, and aims at incorporating several sources of variance, including their interaction with data properties, into an analysis of significance and reliability of machine learning evaluation, with the aim to draw inferences beyond particular instances of trained models. We will show how to incorporate arbitrary sources of noise like meta-parameter variations into statistical significance testing with GLRTs, how to use variance component analysis based on LMEMs to analyze the contribution of noise sources to overall variance, and how to compute a reliability coefficient as indicator for reproducibility.