Multi-Stream Temporal Networks for Emotion Recognition in Children and in the Wild
摘要
In this chapter, we extend and leverage the temporal segment networks framework for emotion recognition in children and in the wild. To that end, we explore the effect of different information streams (Body, Face, Context, Audio, Word Embeddings) and representations (RGB, Flow). We perform an extensive ablation analysis, including the effect of each representation and modality on different emotions, and verify the performance of the proposed systems against the previous SoTA methods in the EmoReact and the BoLD datasets.