错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A dual stream spatio-temporal deep network for micro-expression recognition using upper facial features

  • Nikin Matharaarachchi,
  • Muhammad Fermi Pasha

摘要

The performance of micro-expression recognition mechanisms has greatly improved with the use of deep learning approaches. Most studies have been shifting from the use of hand-designed methods to deep learning-based methods and have increasingly incorporated apex frames into the recognition process. Psychological studies have shown the benefits of only using the upper face for the detection of micro-expressions by humans. However, to the best of our knowledge, this study is the first to explore the effects of using solely upper facial features for computer-based micro-expression detection. To this end, we propose a novel spatio-temporal network to recognize micro-expressions from videos using only upper facial features. The proposed model takes the apex frame and extracts the spatial features using a 2D-CNN and a temporal window of 30 adjacent frames for the temporal feature extraction using a 3D-CNN. The two input streams in the model are in a parallel dual stream architecture. Our proposed method outperforms state-of-the-art methods in detecting micro-expressions in both cropped videos (in the CASMEII and SAMM datasets) and long videos (in the SMIC-E-HS dataset). We further evaluate our proposed method using a combined dataset. Our proposed method achieves a prediction accuracy of 0.832 for the combined dataset, which outperforms state-of-the-art methods. Results show that using only features from the upper face does not impede the detection of micro-expressions and that combining spatial and temporal analysis improves prediction accuracies.