Diagnosing Space and Time Overdispersion Due to Heteroskedasticity
摘要
Introduction: Geotemporal modeling with count-based data requires diagnostic tests to ensure the statistical reliability of results, especially when assumptions like equidispersion and independence of observations are violated. Overdispersion, heteroskedasticity, and spatial or temporal autocorrelation are common in spatial data and require assessment to prevent distortion in model predictions. We demonstrate a methodology for diagnosing spatial and temporal errors using zipcode tabulation area (ZCTA) level data on uninsured populations within a 23-county area in central-southern Florida, focusing on addressing overdispersion, heteroskedasticity, and testing critical model assumptions. Methods: We used five American Community Survey (ACS) 5-year datasets (2014–2022) to diagnose model stability of uninsured populations over time and space. Techniques included multivariate regression, Poisson and negative binomial regression, spatial autocorrelation (Moran’s I), cluster mapping (Getis-Ord Gi*), and GARCH modeling. R and ArcPro 2.6 were employed for analysis. Results: Model diagnostics and automated decision-making based on variable significance and model selection criteria showed systematic assessment of model stability in error-prone secondary data. Unemployment and disability status were consistent predictors across all year groups, while spatial analysis revealed persistent clustering of uninsured populations that decreased over time. GARCH modeling identified regional volatility patterns. Discussion: Spatial models complement non-spatial ones by identifying clusters of uninsured populations, providing valuable context for interpreting predictors’ effects. Even with error-prone data, the utilization of diagnostics in geostatistical modeling can detect overdispersion, spatial dependencies, and temporal heteroskedasticity and allow for making informed decisions about the analysis. These methods address linear, spatial, and temporal errors in geostatistical modeling, improving model stability.