Region attention and label embedding for facial action unit detection
摘要
The detection of Facial Action Units (AUs) is crucial for applications in human–computer interaction and affective computing. However, accurately identifying subtle facial muscle movements and modeling the complex relationships between different AUs remain challenging tasks. To address these challenges, we propose a novel framework comprising two key modules. First, our Region Attention Module focuses on detecting facial muscle movements without relying on any auxiliary information and achieve better AU location. Second, the AU Correlation Learning Module aims to capture intricate AU relationships by leveraging an AU label embedding method. This module represents AU-specific semantics through embeddings and encodes them with facial visual features, effectively enhancing the representation ability of AU features by incorporating learned AU relationships. Our method achieved impressive results, with scores of 65.4% on the BP4D benchmark and 65.5% on the DISFA benchmark.