Non-negative Tensor Representation and Unsupervised Classification of Object Pose in Continuous Image Sequences
摘要
The analysis of video data in machine learning classification tasks has become a significant topic in both research and application areas. The accurate classification of video data by frame sets, particularly when the content contains objects that are in a state of dynamic change, represents a significant and complex undertaking. With regard to the temporal phase of video, the thesis proposes an unsupervised classification method based on non-negative tensor factorization (NTF). In order to transform the video data into low-dimensional data that can be classified by unsupervised clustering, the non-negative tensor factorization method is used to reduce the dimension of the input tensor, resulting in the generation of a non-negative rank-1 tensor combination. Following dimensionality reduction, the extracted features are divided into clusters through Mean Shift, thereby facilitating classification and recognition of video sequences. Experiments on the Coil-100 dataset of item image sequences at different angles and other dynamic pose-changing videos demonstrate the effectiveness of the method.