Gesture-Based Machine Learning for Enhanced Autonomous Driving: A Novel Dataset and System Integration Approach
摘要
This paper describes a new multi-modal dataset for human pose recognition and its use for gesture interaction with autonomous driving vehicles for the purpose of fine positioning them. The dataset consists of 422,036 images grouped into 6 classes representing typical poses for giving instructions to robot vehicles. For each image RGB data, depth data and skeletal data was collected with multi-modal data fusion in mind. It should be emphasised that the entire data set was recorded by a single person, bearing in mind that the combination with depth and skeletal data may mask physical and ethnic characteristics of the subject. For evaluation of the dataset a ResNet101 was used to perform a t-SNE analysis as well as some hidden layers were used from the network to perform a cosine similarity calculation to find duplicates in our dataset.