In order to remain competitive, Internet companies often collect and analyze user data to enhance user experiences. Frequency estimation is a widely used statistical technique, but it can potentially conflict with privacy regulations. To address this, privacy-preserving analytical methods based on differential privacy have been proposed. These methods typically require either a large user base or a trusted server. While these requirements may be feasible for larger companies, they may present challenges for smaller organizations. To overcome this limitation, we introduce a distributed, privacy-preserving, sampling-based frequency estimation method in this chapter. This method maintains high accuracy, even with a small number of users, and does not require a trusted server. The approach combines multi-party computation and sampling techniques to achieve this. Additionally, we establish a relationship between the privacy guarantee, output accuracy, and the number of participants. Unlike most existing methods, our approach offers a centralized differential privacy guarantee without the need for a trusted server. Our results demonstrate that, even with a small number of participants, our method can produce estimates with high accuracy. This provides smaller companies with greater opportunities for growth through privacy-preserving statistical analysis. Furthermore, we propose an architectural model to support weighted aggregation, which improves the accuracy of estimates by accommodating users with varying privacy requirements. Our method outperforms unweighted aggregation, delivering more precise estimates. Extensive experimental results confirm the effectiveness of the proposed methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Differentially Private Distributed Frequency Estimation

  • Xun Yi,
  • Xuechao Yang,
  • Xiaoning Liu,
  • Andrei Kelarev,
  • Kwok-Yan Lam,
  • Mengmeng Yang,
  • Xiangning Wang,
  • Elisa Bertino

摘要

In order to remain competitive, Internet companies often collect and analyze user data to enhance user experiences. Frequency estimation is a widely used statistical technique, but it can potentially conflict with privacy regulations. To address this, privacy-preserving analytical methods based on differential privacy have been proposed. These methods typically require either a large user base or a trusted server. While these requirements may be feasible for larger companies, they may present challenges for smaller organizations. To overcome this limitation, we introduce a distributed, privacy-preserving, sampling-based frequency estimation method in this chapter. This method maintains high accuracy, even with a small number of users, and does not require a trusted server. The approach combines multi-party computation and sampling techniques to achieve this. Additionally, we establish a relationship between the privacy guarantee, output accuracy, and the number of participants. Unlike most existing methods, our approach offers a centralized differential privacy guarantee without the need for a trusted server. Our results demonstrate that, even with a small number of participants, our method can produce estimates with high accuracy. This provides smaller companies with greater opportunities for growth through privacy-preserving statistical analysis. Furthermore, we propose an architectural model to support weighted aggregation, which improves the accuracy of estimates by accommodating users with varying privacy requirements. Our method outperforms unweighted aggregation, delivering more precise estimates. Extensive experimental results confirm the effectiveness of the proposed methods.