Federated Scaling of Pre-trained Models for Deep Facial Expression Recognition
摘要
Building an efficient deep learning-based Facial Expression Recognition (FER) system is challenging due to the requirements of large amounts of personal data and the rise in data privacy concerns. Federated learning has emerged as a promising solution for such problems, which however is communication-inefficient. Recently, pre-trained models have shown effective performance in federated learning setups regarding convergence. In this paper, we extend the traditional FER towards a new paradigm, where we study the performance of federated fine-tuning of standard vision pre-trained models for FER. More specifically, we propose a Federated Deep Facial Expression Recognition (FedFER) framework, where clients jointly learn to fuse the representations generated by pre-trained deep learning models rather than training a large-scale model from scratch without sharing any data. With the help of extensive experimentation using standard pre-trained vision models (ResNet-50, VGG-16, Xception, Vision Transformers) and benchmark datasets (CK+, FERG, FER-2013, JAFFE, MUG), this paper presents interesting perspectives for future research in the direction of federated Deep FER.