Enhanced modeling approaches for count data analysis with focus on substance use outcomes
摘要
The selection of appropriate statistical models is essential for accurately interpreting the analysis of count data, especially in behavioral medicine. Traditionally, Poisson and Negative Binomial models have been commonly employed, but they may not always be the most optimal choices, particularly when dealing with data with an abundance of zeroes, which can be effectively modeled using zero-inflated and zero-altered (hurdle) models. Additionally, U-shaped distributions where the data are clustered around both ends—low and high counts—with fewer occurrences in the middle, cannot be adequately captured by traditional approaches and further complicate the analysis. This paper critically examines the widespread use of zero-inflated Poisson (ZIP) and zero-inflated negative binomial (ZINB) models in the context of adolescent substance use data, identifying their potential limitations. Using a dataset from a smoking study of 1263 adolescents who reported smoking behavior across eight waves, we analyzed the sparse count outcome "Days Smoked in the Past Month," with covariates such as sex, age, and GPA recorded at each wave. Through a comprehensive evaluation of smoking behavior count outcomes—employing model identification via the Kolmogorov–Smirnov (KS) test, validation through confirmation studies, and regression analysis guided by Akaike Information Criterion (AIC). The range of models covered includes: ZIP, Poisson hurdle (PH), ZINB, negative binomial hurdle (NBH), zero-inflated negative binomial with fixed r (ZINB-r), negative binomial hurdle with fixed r (NBH-r), zero-inflated beta-binomial (ZIBB), beta-binomial hurdle (BBH), zero-inflated beta-binomial with fixed n (ZIBB-n), beta-binomial hurdle with fixed n (BBH-n), zero-inflated beta-binomial with fixed alpha and beta (ZIBB-ab), beta-binomial hurdle with fixed alpha and beta (BBH-ab), zero-inflated beta negative binomial (ZIBNB), and beta negative binomial hurdle (BNBH). Our study demonstrates the superior model fitting and regression analysis capabilities of the ZIBB and BBH models. Notably, our findings reveal the effectiveness of the ZIBB model in capturing the U-shaped distribution observed in real-world data. This underscores the importance of exploring a wider range of models beyond ZIP and ZINB for count data analysis. This study advocates for the broader application of these more sophisticated models in behavioral medicine, with the goal to enhance the accuracy and reliability of research outcomes.