Privacy-preserving crowdsourcing labeling with quality control: a differential privacy approach
摘要
Crowdsourcing labeling plays an indispensable role in modern data-driven applications, as it efficiently assists in tackling large-scale labeling tasks. To ensure the utility of data, recent works of crowdsourcing labeling deploy various quality control strategies. However, these strategies have exposed crowdsourcing workers to privacy threats. Current efforts on privacy protection in crowdsourcing mostly focus on addressing privacy concerns in spatial crowdsourcing rather than crowdsourcing labeling. Moreover, they overlook the issue of balancing privacy and data quality. In light of this, we propose a novel differential privacy crowdsourcing labeling algorithm with quality control to protect worker privacy. In this work, we incorporate workers’ confidence into the quality control strategy and provide a differentially private selection mechanism to select capable workers. This mechanism can strictly preserve privacy while maintaining sufficient data utility. To further ensure label quality, we design a consensus-based label aggregation algorithm that filters out unreliable answers, thus mitigating the impact of noise introduced by the privacy mechanism on data utility. The proposed algorithm reduces the dependence of the existing quality control strategy on historical information and realizes worker-level privacy protection through the randomness introduced by differential privacy. Experimental results on six real-world datasets demonstrate that our algorithm achieves a good balance between high-quality labels and strict privacy protection.