Which geographical scale best aligns with epidemiological data?
摘要
Datasets often include geographical locations, making it essential to effectively capture the spatial relationships within such data. In particular, geographical data play a crucial role in studying the spread of epidemics across regions. Applying epidemiological models that account for spatial heterogeneity requires partitioning the study area into meaningful regions. This partition must be appropriately scaled—small enough to capture local dynamics yet large enough to ensure a sufficient population. Additionally, the number of regions should remain manageable to facilitate analysis. To generate partitions with any desired number of regions, we propose a community detection method for networks weighted by a relevant variable. Our approach leverages generalized modularity matrices and leading eigenvectors to create an initial partition. We include two algorithms to refine the partition by identifying sub-communities or forming supra-communities, enabling flexible adjustments to the desired number of regions. Each detected community corresponds to a geographically connected region. A key advantage of our methodology is its ability to capture intra-community heterogeneity by assigning a level of membership to each node, while also recognizing the hierarchical relevance of their connections. We constructed networks for the states of Guanajuato and Jalisco in Mexico, weighted using COVID-19 incidence data. Our method outperforms Leiden, Louvain, Combo, and two variants of Spectral Clustering. For both networks, we identified key nodes within each community based on the level of membership assigned to the municipalities. The geographical regions identified by our method in each state closely align with the official administrative regions defined by the respective governments.