TRAE: Reversible Adversarial Example with Traceability
摘要
Users often upload images containing sensitive personal information to social media platforms. Unfortunately, these images are susceptible to misuse by malicious entities, particularly for training deep neural networks. To protect privacy, current efforts are concentrated on utilizing reversible adversarial examples (RAE) to confuse and disrupt the training of deep neural networks, allowing authorized users to recover the original images. However, existing RAE methods lack research on traceability, leaving protected images vulnerable to attacks, especially when redistributed. To bolster image protection, we propose a secure solution called TRAE. TRAE provides a unified framework for both image traceability and reversible adversarial protection. It is built upon two sets of encoder-decoder pairs: two encoders for generating perturbations and embedding watermarks, and two decoders for extracting watermarks and restoring images. Experimental results demonstrate that images generated by TRAE exhibit favorable visual quality, strong attack capabilities, and efficient restoration capabilities across various datasets. Furthermore, the extracted watermark information maintains a high level of integrity.