Using Clustering Methods to Enhance Data Representativeness: Toward a Well-Being Indicator for Corsican Municipalities
摘要
Are indicators based primarily on economic data sufficient to represent a region’s quality of life? This study aims to develop a well-being indicator for Corsican municipalities to assist decision-makers. Traditional indices like Gross Domestic Product (GDP) and the Human Development Index (HDI) have limitations, often overlooking regional factors, as quality of life is deeply influenced by local territory, social context, and cultural background. Thus, developing indicators at the municipal level is crucial to better reflect local conditions and support decision-making in smaller communities. In this study, we use machine learning to enhance data collection and apply clustering methods to group municipalities with similar characteristics, thereby optimizing sampling efforts. We compare three popular clustering algorithms: Affinity Propagation, K-means, and Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN). Our approach reduces the 360 Corsican municipalities to four distinct groups, each sharing key quality-of-life attributes. We discuss our data collection process, the performance of the clustering algorithms, and potential future research directions.