A Practical Privacy Protection Procedure for Gathering Categorical Data
摘要
Randomized response (RR) is a primary tool for protecting respondent’s privacy in collecting data on categorical variables and its design depends crucially on the privacy criterion. We discuss certain shortcoming of some existing privacy criteria. We show that an average security criterion may not give any privacy protection for some responses. We also find that several other criteria, which simply impose upper bounds on the parity of the RR design, inflict severe data utility loss, unless the number of categories is fairly small. This applies to local differential privacy (LDP), which is a leading privacy criterion. We reveal substantial statistical inefficiency of the RAPPOR procedure, which has been in use by Google, Apple and others. We propose a new privacy procedure that is similar to l-diversity but, works locally for each respondent. The procedure is simple to implement and its privacy promise is easy to understand and communicate to survey participants. We give an unbiased estimator of the probability vector of all categories and prove its minimaxity within a class of estimators under squared error loss. When the number of categories is moderately large, our procedure gives a better privacy-utility trade-off than LDP.