KFM-SLR: A Novel Privacy-Aware Framework for Sign Language Recognition Using Keypoint-Filled Face Masking
摘要
Sign Language Recognition (SLR) has made significant strides in improving accuracy and efficiency. However, the integration of facial data to capture non-manual signals raises critical privacy concerns, particularly in sensitive applications. To address this, we propose Keypoint-Filled Face Masking for SLR (KFM-SLR), a privacy-aware framework that replaces raw facial images with keypoint-based representations. This approach preserves essential structural information while ensuring robust privacy protection. A key innovation of KFM-SLR is its use of self-supervised learning (SSL), which leverages vast amounts of unlabeled sign language data (often containing private facial information) to pretrain models without manual annotation. This is particularly feasible due to the abundance of unannotated yet privacy-sensitive SLR datasets. Subsequently, the model can be fine-tuned using generic, non-private sign language datasets, ensuring strong performance without compromising user privacy. The framework employs a two-stream architecture with three heads design to extract multi-dimensional sign language features, enhancing recognition accuracy. Experiments on German sign language datasets (Phoenix-2014 and Phoenix-2014T) demonstrate that KFM-SLR achieves state-of-the-art (SOTA) recognition performance while maintaining stringent privacy safeguards. These results highlight the framework’s potential for privacy sensitive SLR applications, offering a solution to reconcile accuracy and ethical data usage.