CLIP-AL: an adaptive CLIP-based active learning framework for offline signature forgery verification
摘要
Signature verification is a cornerstone of biometric authentication, where document authenticity and forgery prevention are crucial. This study introduces an improved signature authentication framework that combines Contrastive Language–Image Pretraining (CLIP) embeddings with an active learning loop to reliably detect forged and synthetic signatures. The CLIP model extracts robust multimodal representations of signatures, while the active learning component dynamically queries uncertain predictions for human validation, allowing iterative improvement with limited labeled data. Comprehensive experimental evaluation includes pairwise genuine–forged and genuine–synthetic verification settings, person-disjoint testing, and analysis of uncertainty-driven decision behavior. Experiments on the ICDAR 2011 dataset show notable improvements in accuracy (83.9%), recall (95.9%), and AUC (0.89), outperforming state-of-the-art baselines such as Vision Transformer (ViT), Swin Transformer, and ResNet-50, as well as a CLIP-last static baseline. Additional analyses examine the confidence and calibration characteristics of static baselines to motivate the proposed uncertainty-aware human-in-the-loop design. The proposed method reduces annotation requirements and enhances robustness against synthetic forgeries. These results highlight the promise of integrating CLIP embeddings with human-in-the-loop learning to develop scalable, explainable biometric verification systems for real-world use.