Beyond RGB: Tri-Modal Microexpression Recognition with RGB, Thermal, and Event Data
摘要
Facial Emotion recognition (FER) is an extensively studied computer vision task that aims at identifying and categorizing emotional expressions depicted on a human face, such as anger, fear, or happiness. Due to the subjective nature of feelings, deep learning models may struggle to learn implicit information about a person’s emotions, leading to inaccuracies in existing methods. In this work, we aim to estimate microexpressions-small facial movements that can indicate underlying feelings, as described in the Facial Action Coding System (FACS)-from face videos, as these facial movements provide explicit information that is more easily perceivable by deep learning architectures. Furthermore, despite the evolution of FER technologies driven by advancements in neural network architectures and the exploration of new sensing technologies, there is a significant shortage of datasets that leverage these emerging modalities, which limits the progress of research in this field. In our study, we aim to explore and compare the feasibility of using different input data modalities, visible, thermal, and event, as training and testing data for a CNN baseline network by presenting a pioneering dataset that integrates these three modalities, each annotated with detailed Facial Action Units (FAUs) present in the FACS. Our proposed Visible, Event, and Thermal Face Dataset for Micro Expression Recognition (VETEX) containing 2506 face videos is available upon request.