Significance
摘要
Unfortunately, in state-of-the-art research following the train-dev-test paradigm, systematic uncertainty estimation is a neglected problem, and statistical significance testing is often completely ignored. The goal of this chapter is to promote model-based significance testing using LMEMs, and to revitalize the generalized likelihood ratio test (GLRT) as a hypothesis test framework that shows its full potential in a model-based setting. GLRTs date back to the famous Neyman–Pearson theory of statistical testing, and provide a general hypothesis testing framework that applies to any evaluation metric, to multiple meta-parameter settings, and allows analyzing performance differences conditional on test data properties.