Self-supervised rotation-aware guided attentive distilled transformer network for crowd face detection
摘要
Accurate crowd face detection is vital in surveillance, security, and crowd management due to challenges like varying densities, occlusions, and scale variations. Resizing issues arise when traditional techniques fail to handle diverse face sizes and densities, alongside patterns or objects resembling faces, hindering accuracy. To address this, we propose the Self-Supervised Rotation-aware Guided Attentive Distilled Transformer Network (SSRGADTN) for crowd face detection. This method combines self-supervised learning and a rotation-aware distilled transformer network, effectively utilizing unlabeled data to learn robust representations and handle diverse crowd scenes. The workflow starts with self-supervised training on the FWL dataset, followed by integration with the Guided Attentive Detector Network (GADN) trained on the WIDERFACE dataset to enhance face detection, forming the SSRGADN (Self-Supervised Rotation-aware Guided Attentive Detector Network) model. Knowledge distillation further refines the model using a transformer network as the student model trained on the CrowdHuman dataset, resulting in the SSRGADTN model, which outperforms other state-of-the-art models in crowd face detection accuracy.