ViT-LPATA: a vision transformer model for autism detection in children using facial images
摘要
To address the difficulty in recognizing subtle differences in facial biomarkers in children with autism, a learnable positional encoding enhancement (LPEE) module was combined with the adaptive token aggregation (ATA) module. The vision transformer with learnable positional encoding and adaptive token aggregation (ViT-LPATA), a predictive model for autism, was proposed. The model leverages the LPEE module to dynamically capture facial geometric deformation features and integrates the ATA module to enhance the feature representation capability of pathological regions, thereby establishing precise mappings of biomarker differences. Experiments on a publicly available autism facial dataset demonstrated that the ViT-LPATA achieved optimal performance, with 99.2% accuracy and an area under the curve (AUC) value of 0.940.