In today’s interconnected and dynamic global landscape, addressing key challenges such as public health, environmental sustainability, and social inequality demands coordinated, data-informed policy responses. Addressing these issues requires coordinated and evidence-based policy interventions, for which national statistical offices (NSOs) play a crucial role by providing high-quality and timely data. However, the presence of missing or incomplete information—whether arising from surveys or administrative sources—represents a persistent threat to the validity and reliability of statistical outputs. This study investigates the potential of advanced imputation techniques based on supervised machine learning (ML) and deep learning (DL) approaches, which fall within the broader domain of Artificial Intelligence (AI), to improve the handling of missing data in official statistics. A comparative analysis is conducted using real-world microdata from the Italian National Institute of Statistics (Istat), specifically from the Multipurpose Survey on Households, which is characterized by a non-negligible incidence of item nonresponse. Preliminary findings indicate that AI-driven imputation strategies outperform traditional statistical methods in terms of accuracy and robustness, particularly in complex social datasets. The results contribute to the growing body of literature advocating for the integration of modern computational tools within the framework of official statistics, with the aim of enhancing data quality on critical societal issues.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Innovative Applications of Supervised Learning in Addressing Missing Data: A Case Study on Social Surveys

  • Simona Cafieri,
  • Francesco Pugliese,
  • Mauro Sodani

摘要

In today’s interconnected and dynamic global landscape, addressing key challenges such as public health, environmental sustainability, and social inequality demands coordinated, data-informed policy responses. Addressing these issues requires coordinated and evidence-based policy interventions, for which national statistical offices (NSOs) play a crucial role by providing high-quality and timely data. However, the presence of missing or incomplete information—whether arising from surveys or administrative sources—represents a persistent threat to the validity and reliability of statistical outputs. This study investigates the potential of advanced imputation techniques based on supervised machine learning (ML) and deep learning (DL) approaches, which fall within the broader domain of Artificial Intelligence (AI), to improve the handling of missing data in official statistics. A comparative analysis is conducted using real-world microdata from the Italian National Institute of Statistics (Istat), specifically from the Multipurpose Survey on Households, which is characterized by a non-negligible incidence of item nonresponse. Preliminary findings indicate that AI-driven imputation strategies outperform traditional statistical methods in terms of accuracy and robustness, particularly in complex social datasets. The results contribute to the growing body of literature advocating for the integration of modern computational tools within the framework of official statistics, with the aim of enhancing data quality on critical societal issues.