TDZS: top semantic embedding and dynamic feature matching for zero-shot skeleton action recognition
摘要
Skeleton action recognition has advanced considerably in recent years, but the integration of skeleton action recognition with zero-shot learning remains relatively underexplored. Traditional mainstream methods primarily focus on aligning textual and skeleton data, often neglecting the semantic feature compensation for skeleton visual features, and Semantic feature compensation involves using multiple semantic features to fill the information gaps in skeletal visuals, providing different descriptive references for visual features through rich semantic information. To address this limitation, this paper designed the top semantic embedding (TSE) framework to enrich the embedding of semantic features into visual features. Through a scoring mechanism, the framework selects the best semantic-visual feature pairs to improve the model’s learning performance. Additionally, to further strengthen the connection between semantic and visual modalities, the Dynamic Feature Matching (DFM) method was proposed. By constructing a multimodal attention matrix between semantic and visual features, the DFM module enables visual features to adaptively match the most relevant semantic features, creating tighter connections between the two modalities. Experiments were conducted on three benchmark datasets. The proposed method achieved an accuracy of 82.63% on the NTU-60 dataset, 77.23% on the PKU-MMD dataset, and 58.21% on the NTU-120 dataset. The experimental results demonstrate the effectiveness and strong performance of the proposed method, similarly, compared to the state-of-the-art method SMIE, it shows improvements of 4.65%, 8.08%, and 1.11% on the three datasets, respectively.