Income inequality under partial observability: What can we learn from incomplete survey data?
摘要
Partial identification techniques offer a robust methodological framework for addressing missing data in survey research without relying on untestable assumptions about non-response mechanisms. We apply this approach to income inequality measurement using the Chilean National Socioeconomic Characterization Survey (CASEN), demonstrating how conventional imputation methods can systematically bias distributional estimates. Our methodology shows that, despite a seemingly modest nonresponse rate of 2.7% in CASEN, the lack of information on income generates considerable uncertainty in the estimation of distributional features. In particular, we show that even this limited missingness leads to wide bounds for key quantiles of the income distribution and for the conditional distribution of income given educational attainment. Moreover, for the observed mean, the plausible range for the Gini coefficient spans from 0.379 to 0.415, a difference large enough to reverse standard poverty assessments and policy conclusions. Additionally, the official imputation procedures tend to overestimate incomes for the poor and underestimate incomes for the rich, leading to downward-biased inequality measures. Unlike existing correction methods that require strong distributional assumptions or privileged access to administrative data, our partial identification framework offers assumption-free bounds that are broadly applicable across household surveys. We also provide open-source computational tools, making the methodology easily adoptable by statistical offices and researchers worldwide.