A Flexible Keyword-Based PIR Scheme with Customizable Data Scales for Multi-server Learning
摘要
As the digital era progresses, data has become an invaluable asset for businesses and organizations. Conducting in-depth analyses of specific data categories requires a strategy that ensures data privacy while effectively leveraging distributed data resources. However, existing technologies face two major challenges: first, accurately defining the boundary of training data is difficult, leading to the inclusion of irrelevant data that can negatively impact model training accuracy; second, user-side retrieval privacy is vulnerable to breaches, potentially exposing the user’s training intentions to the server. To address these challenges, we propose a novel user-proxy server-multi-server framework within the context of Private Information Retrieval (PIR), designed to protect user privacy while learning from specific datasets. Building on this framework, we introduce a keyword-based PIR protocol tailored for multi-server deep learning models. This protocol allows users to query and retrieve targeted datasets from servers via a proxy server, enabling high-precision model training with specific data subsets. Experimental results demonstrate that, under various query volumes and server configurations, the proposed framework significantly reduces response times and improves query efficiency. Additionally, it exhibits remarkable scalability and adaptability, handling dynamic database updates and varying dataset sizes. This approach provides a secure and effective solution for training deep learning models across multiple data sources, making it particularly suitable for applications in smart factories and intelligent manufacturing environments, where data privacy is critical.