Enhanced Facial Emotion Recognition Using Vision Transformer Models
摘要
Automation of facial emotion recognition is an important branch of artificial intelligence and computer vision that has many potential applications in mental health diagnostics, human–computer interaction and security. The existing methods, however, usually have weaknesses in robustness, scalability and computational efficiency. This work proposes a self-attention-based Vision Transformer method that treats images as sequences of patches to capture global dependencies and spatial relations more effectively than other methods. The model is trained and evaluated using a large-scale dataset. On average, the model achieves an overall accuracy of 97%, with good precision, recall and F1 scores in most emotion categories. The model performed better and was more robust to variations in illumination and facial pose compared to other existing methods. This work takes a step forward in facial emotion recognition technology, providing a large-scale and efficient solution for real-world applications. Facial Emotion Recognition, a New Vision Transformer Based on Self-Attention for Machine Learning.