<p>Reproducibility of research is crucial and has received much attention in recent years. One aspect of reproducibility is statistical reproducibility, which examines whether statistical inferences remain similar when experiments are repeated. This paper investigates the reproducibility probability (RP) of normality tests and tests for equality of variances from a nonparametric predictive inference (NPI) perspective, offering a novel predictive framework for quantifying test reproducibility without strong parametric assumptions. Three well-known normality tests–the Shapiro-Wilk (SW), Anderson-Darling (AD), and Lilliefors (LF) tests–are studied, along with two tests for equality of variances: the F-test and Levene’s test. The results show that RP tends to be low, particularly when p-values are close to the significance threshold. RP is also influenced by sample size and significance level, with larger samples decreasing RP in the non-rejection area and increasing it in the rejection area. Among the normality tests, the Shapiro-Wilk test has the highest RP in the non-rejection area, while the Anderson-Darling test has the highest RP in the rejection area. For the equality of variances tests, the F-test exhibits greater variability, particularly under non-normality. These findings highlight the limited statistical reproducibility of widely used tests and demonstrate how the proposed NPI-based approach can provide practical insight into the stability of test outcomes under uncertainty.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On Statistical Reproducibility of Normality and Equality of Variances Tests

  • Norah D. Alshahrani,
  • Tahani Coolen-Maturi,
  • Frank P. A. Coolen

摘要

Reproducibility of research is crucial and has received much attention in recent years. One aspect of reproducibility is statistical reproducibility, which examines whether statistical inferences remain similar when experiments are repeated. This paper investigates the reproducibility probability (RP) of normality tests and tests for equality of variances from a nonparametric predictive inference (NPI) perspective, offering a novel predictive framework for quantifying test reproducibility without strong parametric assumptions. Three well-known normality tests–the Shapiro-Wilk (SW), Anderson-Darling (AD), and Lilliefors (LF) tests–are studied, along with two tests for equality of variances: the F-test and Levene’s test. The results show that RP tends to be low, particularly when p-values are close to the significance threshold. RP is also influenced by sample size and significance level, with larger samples decreasing RP in the non-rejection area and increasing it in the rejection area. Among the normality tests, the Shapiro-Wilk test has the highest RP in the non-rejection area, while the Anderson-Darling test has the highest RP in the rejection area. For the equality of variances tests, the F-test exhibits greater variability, particularly under non-normality. These findings highlight the limited statistical reproducibility of widely used tests and demonstrate how the proposed NPI-based approach can provide practical insight into the stability of test outcomes under uncertainty.