<p>This paper examines the use of <i>p</i>-values in the null hypothesis significance test and the challenges related to their interpretation. It critiques <i>p</i>-values, particularly in large-scale data analysis, and evaluates the <i>d</i>-value as an alternative. Proposed by Demidenko (Am Stat 70(1):33–38, 2016), the <i>d</i>-value measures the probability that an observation from one group exceeds that from another. The <i>d</i>-value emphasizes the magnitude of practical effects, making it valuable in fields such as social sciences. Although the sample size was confirmed to have a minimal direct impact on the <i>d</i>-value itself, its critical role in the precision and stability of the estimate is underscored. Crucially, this study demonstrates the profound and often counterintuitive influence of differences in means, standard deviations, and particularly data skewness, on <i>d</i>-value. Using variations in both the normal and gamma distributions, this exploration reveals that the <i>d</i>-value is sensitive to deviations from normality, especially in the presence of differential skewness between groups. The <i>d</i>-value was applied to random samples from the 2019 Brazilian higher education microdata, analyzing variables such as gender, shift, and academic degree. An index of study hours completion was created to facilitate comparisons. The results indicated that, despite large samples, performance differences across categories were minimal. The <i>d</i>-value remained stable, unlike <i>p</i>-values.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Beyond significance: why the d-value matters more in big data contexts

  • Jeff Caponero,
  • Lilia Carolina Carneiro da Costa,
  • Anderson Ara

摘要

This paper examines the use of p-values in the null hypothesis significance test and the challenges related to their interpretation. It critiques p-values, particularly in large-scale data analysis, and evaluates the d-value as an alternative. Proposed by Demidenko (Am Stat 70(1):33–38, 2016), the d-value measures the probability that an observation from one group exceeds that from another. The d-value emphasizes the magnitude of practical effects, making it valuable in fields such as social sciences. Although the sample size was confirmed to have a minimal direct impact on the d-value itself, its critical role in the precision and stability of the estimate is underscored. Crucially, this study demonstrates the profound and often counterintuitive influence of differences in means, standard deviations, and particularly data skewness, on d-value. Using variations in both the normal and gamma distributions, this exploration reveals that the d-value is sensitive to deviations from normality, especially in the presence of differential skewness between groups. The d-value was applied to random samples from the 2019 Brazilian higher education microdata, analyzing variables such as gender, shift, and academic degree. An index of study hours completion was created to facilitate comparisons. The results indicated that, despite large samples, performance differences across categories were minimal. The d-value remained stable, unlike p-values.