Action Segmentation Based on Encoder-Decoder and Global Timing Information
摘要
Action segment has made significant progress, but segmenting and recognizing actions from untrimmed long videos remains a challenging problem. Most state-of-the-art (SOTA) methods focus on designing models based on temporal convolution. However, the limitations of modeling long-term temporal dependencies and the inflexibility of temporal convolutions restrict the potential of these models. To address the issue of over-segmentation in existing action segmentation algorithms, which leads to prediction errors and reduced segmentation quality, this paper proposes an action segmentation algorithm based on Encoder-Decoder and global temporal information. The action segmentation algorithm based on Encoder-Decoder and global timing information proposed in this paper uses the global timing information captured by LSTM to assist the Encoder-Decoder structure in judging the action segmentation point more accurately and, at the same time, suppress the excessive segmentation phenomenon caused by the Encoder-Decoder structure. The algorithm proposed in this paper achieves 93% frame accuracy on the constructed real Taiji action data set. The experimental results prove that this model can accurately and efficiently complete the long video action segmentation task.