Pixel-Level Intelligence for Rotation-Invariant Face Detection with Guided Deformable Attention
摘要
Detecting rotated faces has always been a challenging task. Fixed convolutional kernels struggle to effectively match features after rotation, while the sampling point offsets of deformable convolutions are limited by complex backgrounds. To address this issue, we propose a guided deformable attention network. Guiding the offset direction of sampling points by adding constraints of facial structure to deformable convolutions. Our network adopts a dual-stream structure, with one branch detecting the inherent structural information for preliminary positioning of the face area; then, the second branch uses deformable convolution to perform pixel-level feature extraction on the face within the range. In addition, we introduce a novel loss, which, during the guidance process, aligns the activation areas in the feature maps extracted by the two branches through the Kullback–Leibler divergence. Extensive experimental results validate that our proposed network performs excellently on multiple face detection datasets, surpassing the current state-of-the-art face detection methods.