Enhancing Privacy Preservation in Speech Data Publishing
摘要
In speech data publishing, users’ data privacy is disclosed, and thereby more privacy of users is breached, since speech data contains a large amount of information about speakers. Existing work focused on sanitization in speech content, speakers’ voice, and data descriptions, without considering the correlation of speech content and speaker’s voice. Therefore, these existing work cannot protect speakers’ data privacy when attackers utilize such correlation to identify speakers’ speech data. To tackle this problem, in this paper, we propose a protocol to decrease such potential risk in speech data publishing while keeping the balance of privacy preservation and data utility. Specifically, we define both the risks of privacy disclosure and the data utility loss in speech content, speaker’s voice, and dataset description. Moreover, we do the first attempt to formalize the correlation between speech content and speaker’s voice, and regard it as a new kind of privacy leakage risk. Thereafter, we utilize the classifier in Machine Learning and optimize the speech data sanitization considering the defined risks of privacy disclosure and data utility loss. Finally, simulation results validate the effectiveness of the proposed protocol.