FMCBGait: Fine-grained multimodal gait recognition with cues embedding and body parts distribution guidance
摘要
Low information entropy and coarse-grained motion pattern descriptions remain critical challenges for achieving accurate gait recognition in complex environments. To address these limitations, we propose FMCBGait, a novel fine-grained multimodal gait recognition framework that leverages the complementary strengths of gait parsing and skeleton modalities. In the parsing modality, we employ fine-grained segmentation of human semantic structures to extract body part patches of varying sizes, which are subsequently normalized to a unified scale using the Multi-Scale Adaptive Patch Embedding (MSAPE) module for effective feature encoding. For the skeleton modality, we construct a skeleton hypergraph to represent human topology, enabling the capture of higher-order interactions through a Hypergraph Convolutional Network (HGCN). To enhance inter-modal complementarity, representative visual cues from parsing data are embedded into the skeleton hypergraph, while structural cues from the skeleton modality are reciprocally embedded into the parsing modality. Additionally, the Multi-Stage Modal Fusion Strategy (MSMFS) integrates these modalities through early-stage cues embedding, mid-stage Body Part Distribution Guidance (BPDG), and late-stage Heterogeneous Non-Local (HNL) attention, facilitating comprehensive feature fusion. FMCBGait achieves state-of-the-art performance, with Rank-1 accuracies of 75.5% and 80.7% on two challenging datasets, demonstrating its robustness and efficacy for multimodal gait recognition.