Distributions and Their Approximations for p-Values
摘要
In three examples, we illustrate the distribution of p-values arising in studies involving testing multiple hypotheses. These examples closely follow the one-parameter power distribution but sometimes appear to underestimate the likelihood of observing the very smallest p-values. This behavior also occurs when the underlying test statistic behaves as a Chi-squared or as a two-tailed normal or Student’s t-distribution. We demonstrate that such behavior can appear as a result of either sampling from a mixture of different alternative hypotheses or perhaps due to a lack of independence among the p-values. This motivates our development of a model of dependent sampling. The approximation of the distribution of p-values is useful in explaining the number of identified hypotheses when correcting for multiplicity using any of several popular criteria. We use the order statistics to estimate the number of statistically significant hypotheses in each example and compare these to the observed values. In the simulation study, we compute the average number of statistically significant p-values over iterations and compare with the estimation using the order statistics. The results show that the power distribution is a good approximation of the alternative distribution of p-values, especially for the region corresponding to the statistically significant hypotheses away from the boundary of the extreme p-values.