SILK-SVM: An Effective Machine Learning Based Key-Frame Extraction Approach for Dynamic Hand Gesture Recognition
摘要
Dynamic hand gesture recognition is a fundamental domain in Human–Computer Interaction, wherein machine learning and deep learning models are widely employed. Key-frame extraction from the input hand gesture videos is a challenging task in this domain, since voluminous video datasets contain numerous redundant frames leading to storage inefficiency and longer training times. The existing approaches address this problem using various traditional and deep learning methods for this problem, aimed at extracting a fixed number of key-frames from all input videos. This paper introduces a novel skeleton-based silhouette-optimized K-means clustering mechanism (SILK) for dynamic Key-Frame Extraction. Initially, the hand skeleton features are localized on the video frames followed by frame filtering, consecutively, an unsupervised K-means clustering optimized by Silhouette score metric dynamically extracts an optimal number of unique key-frames from each video based on frame similarity. Effectiveness of the so extracted key-frames has been verified by SVM classification. The proposed approach (named as SILK-SVM) has been thoroughly tested on two publicly available datasets