This document describes an agile strategy in statistical analysis to generate synthetic data to overcome increasing obstacles to carry out face-to-face surveys in Mexico, such as increasing insecurity, limited access to certain areas controlled by organized crime and budgetary constraints. We use two data sources: the 2020 Income-Expenditure Survey or ENIGH by its acronym in Spanish, and the 2020 Population and Housing Census, or CPyV, both carried out by the National Institute of Statistics and Geography of Mexico (INEGI), and several statistical learning techniques such as PCA, clustering, random forest and classification methods to generate granular synthetic data with scientific, policy and commercial uses. The result is an algorithm that allows characterizing the socioeconomic level and the income and expenditure profiles of urban households at block level for all the country, and we suggest metrics to validate the synthetic data due to the impossibility of having disaggregated data to validate our results. This agile strategy can be replicated for various contexts where data layers that satisfy certain basic conditions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mi Casa no Es Tu Casa: An Agile Strategy to Generate Synthetic Data to Overcome Security Challenges in Mexico

  • José Carlos Rodríguez Pueblita,
  • Edgar Oswaldo Díaz

摘要

This document describes an agile strategy in statistical analysis to generate synthetic data to overcome increasing obstacles to carry out face-to-face surveys in Mexico, such as increasing insecurity, limited access to certain areas controlled by organized crime and budgetary constraints. We use two data sources: the 2020 Income-Expenditure Survey or ENIGH by its acronym in Spanish, and the 2020 Population and Housing Census, or CPyV, both carried out by the National Institute of Statistics and Geography of Mexico (INEGI), and several statistical learning techniques such as PCA, clustering, random forest and classification methods to generate granular synthetic data with scientific, policy and commercial uses. The result is an algorithm that allows characterizing the socioeconomic level and the income and expenditure profiles of urban households at block level for all the country, and we suggest metrics to validate the synthetic data due to the impossibility of having disaggregated data to validate our results. This agile strategy can be replicated for various contexts where data layers that satisfy certain basic conditions.