Since GCN has been proposed to represent skeleton data as graphs, it has always been the primary method for skeleton-based human action recognition. However, when dealing with interaction skeleton sequences, current GCN-based methods do not consider dynamically updating the connections between the skeleton points of two persons and cannot extract interaction features well. The self-attention module of Transformer can well focus on the correlation between skeleton sequences. We propose a novel method called Dynamic Skeleton Association Transformer (DSAT) for dyadic interaction action recognition, which can dynamically update the interaction relationship adjacency matrix by combining the spatial attention features and geometric spatial distances of two skeleton sequences to capture the spatial interaction relationship between the skeleton sequence of the two persons. Then, we use spatial self-attention to extract the interaction relationships between different individuals and within the same individual. We also improve the temporal self-attention module according to the density of interactive events to extract the correlation between the same skeleton point in different frames. Through our strategy, our model can more effectively recognize interactive behaviors that are density in time and space, and we have conducted extensive experiments on the benchmark datasets of SBU, NTU-RGB+D, and NTU-RGB+D 120 interaction subsets to verify the effectiveness of our method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dynamic Skeleton Association Transformer for Dyadic Interaction Action Recognition

  • Zixian Liu,
  • Longfei Zhang,
  • Xiaokun Zhao,
  • Yixuan Wang

摘要

Since GCN has been proposed to represent skeleton data as graphs, it has always been the primary method for skeleton-based human action recognition. However, when dealing with interaction skeleton sequences, current GCN-based methods do not consider dynamically updating the connections between the skeleton points of two persons and cannot extract interaction features well. The self-attention module of Transformer can well focus on the correlation between skeleton sequences. We propose a novel method called Dynamic Skeleton Association Transformer (DSAT) for dyadic interaction action recognition, which can dynamically update the interaction relationship adjacency matrix by combining the spatial attention features and geometric spatial distances of two skeleton sequences to capture the spatial interaction relationship between the skeleton sequence of the two persons. Then, we use spatial self-attention to extract the interaction relationships between different individuals and within the same individual. We also improve the temporal self-attention module according to the density of interactive events to extract the correlation between the same skeleton point in different frames. Through our strategy, our model can more effectively recognize interactive behaviors that are density in time and space, and we have conducted extensive experiments on the benchmark datasets of SBU, NTU-RGB+D, and NTU-RGB+D 120 interaction subsets to verify the effectiveness of our method.