A Hybrid Residual and Capsule Layer Based CNN Model for Yoga Pose Estimation
摘要
The escalating global popularity of yoga, known for its multifaceted benefits, has stimulated interest in automated yoga pose estimation—a process that involves identifying and classifying yoga poses from visual data such as photos or videos. This technology bears substantial potential for diverse applications, including personalized feedback to practitioners, progress tracking, and sequence customization. In this paper an effective automated yoga pose estimation framework is proposed by introducing a new hybrid Convolutional Neural Network (CNN) that incorporates residual and capsule layers. Here the residual layer addresses the vanishing gradient problem, while capsule layers capture the hierarchical structure inherent in images, and consequently enhancing feature transfer and utilization efficiency. The proposed model framework discovers a pattern for yoga poses at varying camera distances and also captures the spatial relation between the instances of the images. It is noteworthy for yoga pose estimation, as yoga images frequently contain noise and deformations caused by variations in body type, camera distance, and capturing angle of the practitioners, as well as distinct poses. To comprehensively assess the model’s performance, classification reports, accuracy and loss graphs, and ROC curves are meticulously analysed and presented. By employing multiple optimizers for comparison, it is observed that Adam optimizer yields the highest accuracy at 96.62% with lowest computational complexity. The proposed hybrid model, featuring residual and capsule layers, exhibits promising outcomes in the accurate classification of yoga poses.