Most privacy-preserving action recognition (PPAR) methods primarily address spatial-domain privacy removal, often neglecting privacy risks in the frequency domain. Additionally, current architectures often struggle to capture fine-grained privacy features and handle high-frequency information. To address these issues, we propose a dual-domain multiscale perturbation framework based on adversarial training, combining frequency-domain perturbation and spatial-domain feature learning. The anonymization module utilizes a Swinv2-Unet architecture, incorporating Swinv2-T as the encoder and U-Net as the decoder. A novel Wavelet Frequency Intervention Module (WFIM) decomposes video frames into high-frequency subbands containing privacy-sensitive details and low-frequency subbands conveying action trends. Learnable Laplacian noise suppresses high-frequency privacy information, while attention mechanisms enhance low-frequency action features, embedding these refined features at multiple decoder levels. Additionally, we design a lightweight cross-domain interaction module that dynamically fuses frequency-domain and spatial-domain features using Neighborhood Attention and its cross-attention variant. Experimental results show that our method achieves stronger privacy protection than previous approaches, with only a slight drop in action recognition performance, demonstrating an effective balance between privacy preservation and task utility.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Frequency Perturbation and Spatial Attention Modulation for Privacy-Preserving Action Recognition

  • Jiahui Ding,
  • Xingyuan Chen,
  • Huahu Xu

摘要

Most privacy-preserving action recognition (PPAR) methods primarily address spatial-domain privacy removal, often neglecting privacy risks in the frequency domain. Additionally, current architectures often struggle to capture fine-grained privacy features and handle high-frequency information. To address these issues, we propose a dual-domain multiscale perturbation framework based on adversarial training, combining frequency-domain perturbation and spatial-domain feature learning. The anonymization module utilizes a Swinv2-Unet architecture, incorporating Swinv2-T as the encoder and U-Net as the decoder. A novel Wavelet Frequency Intervention Module (WFIM) decomposes video frames into high-frequency subbands containing privacy-sensitive details and low-frequency subbands conveying action trends. Learnable Laplacian noise suppresses high-frequency privacy information, while attention mechanisms enhance low-frequency action features, embedding these refined features at multiple decoder levels. Additionally, we design a lightweight cross-domain interaction module that dynamically fuses frequency-domain and spatial-domain features using Neighborhood Attention and its cross-attention variant. Experimental results show that our method achieves stronger privacy protection than previous approaches, with only a slight drop in action recognition performance, demonstrating an effective balance between privacy preservation and task utility.