WFocusedGait: wavelet-inspired focused multimodal feature fusion for gait recognition
摘要
As a prominent biometric modality, gait recognition has garnered substantial attention in recent years. Compared to single-modal approaches, multimodal methods offer significant advantages by leveraging multiple sources of information. However, some multimodal methods ignore the frequency information and underutilize spatial–temporal interaction features such as gait movement pattern changes and pose and structure information. To address these limitations, we propose a multi-stage focused feature fusion algorithm based on learnable wavelets, which effectively integrates different spatial–temporal features during the feature extraction process. First, we introduce a gait wavelet convolution that combines the physical frequency features of silhouettes, transforming gait information into time–frequency representations to mitigate the impact of confounding factors such as appearance. Considering the semantic relationship between silhouettes and skeletons, as well as the varying contributions of different information within the gait cycle to recognition, we propose a focal fusion module. Specifically, this module emphasizes globally critical features of similar silhouettes and skeletons and explores diverse spatial–temporal joint features through depthwise separable convolution, thereby expanding the fusion receptive field. Additionally, to address spatial–temporal feature fusion at local and part, we introduce an internal attention module to further integrate information of silhouettes and skeletons. This module extracts robust features from both temporal and spatial perspectives, localizes regions of interest, and enhances the generalization ability of gait information. By integrating the above strategies and modules, our proposed wavelet-focused gait recognition network demonstrates superior performance on two benchmark datasets.