错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Perturbation Methods: An Application to a Real Dataset

  • Jorge Morais,
  • Rita Sousa,
  • Susana Faria

摘要

The demand for data access has been growing a lot in recent years. The compromise between the utility of the information provided and the protection of confidentiality is increasingly important. Statistical disclosure control (SDC) techniques suggest methods for modifying data so that they can be published without revealing confidential information that can be linked to specific individuals (Benschop et al., Statistical Disclosure Control: A Practice Guide. The World Bank (2021); Matthias et a., Package ‘sd-cMicro’. Technical report (2021)). In this study, we describe and compare different perturbation approaches based mainly on linear and nonlinear models. We show the advantages and disadvantages of each method for data perturbation. We also present several measures for data utility and disclosure risk to evaluate the methods’ performance. This chapter illustrates an application of these perturbation methods using functions from sdcMicro package in R (Patrick and Norman, ACM Trans. Database Syst. 19, 47–63 (1994)). In the literature, it is clear that exact general additive data perturbation (EGADP) and data shuffling produce the lowest disclosure risk and the highest data utility (Matthias, Statistical Disclosure Control for Microdata: Methods and Applications in R. English, 1st edn., vol. 1. Springer International Publishing (2017)). However, our conclusion is that the noise models outperformed the best models in the literature for the real microdataset described in (Banco de Portugal Microdata Research Laboratory (BPLIM): Incentives Systems Data. Banco de Portugal (2021). http://doi.org/10.17900/SI.APR2021.V1 ).