Region- and feature-informed inference engine (REFINE) with CoAtNet-0 for facial age estimation
摘要
Despite recent advances in deep vision models, estimating chronological age from unconstrained facial images remains challenging: age cues are sparse and localized, whereas real-world images vary widely in pose, illumination, and quality. Although deep convolutional networks have long been the workhorse for this task, recent vision transformers improve global reasoning but tend to diffuse attention over many low–value regions. To address this gap, we propose a new processing methodology and develop a novel inference engine, hereafter referred to as