Pose pattern mining using transformer for motion classification
摘要
Capitalizing on the rapid development of diverse deep learning technologies in the field of image analysis, studies are now being conducted to detect objects within images, estimate the poses of target objects, and classify motions. However, the magnitude of image data and computational complexity present challenges in performing real-time image analysis. In addition, the classification of human motions, specifically, requires an effective methodology based on analysis of the changing features of poses from frame to frame. To address this, pose pattern mining using a transformer for motion classification is proposed, which expands the output of the neural network of an object detection model into a time-series pose classification model combining EfficientNet and transformer mechanisms. With regard to the structure of the configured model, the object detection model maintains the mechanisms and displays the effect of expanding the output node of the internal neural network into a time-series-based pose-motion analysis model. Sequence pattern mining is then applied to ensure the efficiency of the data analysis. This hybrid methodology achieved a response time of 0.059 s per frame and an accuracy of 86.67%. Therefore, it may be surmised that this proposed method can be applied to fields such as security and surveillance systems that require fast processing times and high levels of accuracy.