Enhancing Privacy and Preserving Accuracy in Medical Image Classification with Limited Labeled Samples
摘要
For large-scale medical data, annotation requires certain medical knowledge and experience, and manual annotation takes a lot of time and human resources. The application of deep learning technology in medical images can help doctors make diagnoses more quickly and accurately, which promotes technological progress in the medical field. However, images may contain sensitive information such as disease details and individual body structures. It has been shown that an attacker can determine whether a patient participates in training by launching a membership inference attack. To prevent the leakage of patient training samples, this paper proposes a privacy-preserving scheme named PATE-Medical for training high-performance medical image classifiers using a small number of labeled samples. It is based on the idea of employing a faculty-student architecture to preserve the privacy of the training data. To overcome the low accuracy and privacy challenges posed by limited medical images, firstly, a Siamese neural network is utilized to train a single faculty model to obtain accurate prediction results with insufficient training samples. Then, the student model is trained as the output classifier model based on the semi-supervised Mixmatch method, which aims to ensure the performance of the student model while reducing the loss of privacy spent on student training. Experiments show that the accuracy of faculty ensemble voting in PATE-Medical reaches 98.25%; the student model achieves a maximum accuracy of 97%. The minimum privacy cost is only 1.154.