错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detecting statistical anomalies in COVID-19 surveillance database

  • Inma Conde,
  • Esther-Lydia Silva-Ramírez,
  • M. Belén Gómez-Tovar,
  • Angel L. Ortiz,
  • F. L. Cumbrera

摘要

The unexpected unfolding of the COVID-19 pandemic involved the rapid gathering, integration, processing, and dissemination of large volumes of epidemiological data worldwide. In this context, heterogeneous reporting practices and data-integration issues may introduce certain statistical anomalies that can affect the scientific reliability of downstream analyses. Because global surveillance systems continuously aggregate multi-country time series, anomaly screening also becomes a computational problem requiring automated, scalable, and potentially near-real-time analytical workflows, which is relevant to large-scale data processing and high-performance computing environments. The purpose of the present study is to perform a quantitative analysis of COVID-19 surveillance data using Newcomb–Benford’s law and calibrated supervised datasets. As a benchmark, daily new deaths from ten countries, two from each continent, were studied, with a total of 738 records per country. The proposed framework moves beyond a binary conforming/non-conforming diagnosis by estimating the magnitude of deviation through calibrated descriptors and principal-component-based quantification. The comparison between High Income and Low Income aggregates yielded anomaly estimates of 15.27% and 20.55%, respectively, with an uncertainty of 5.2%. Overall, the results indicate that several datasets exhibit non-negligible statistical deviations from the expected digit distribution, interpreted simply as anomaly signals that may support large-scale epidemiological data validation.