Single-Stage Dual-Task Joint Learning Framework for Hand Hygiene Assessment
摘要
In recent years, Action Quality Assessment technology has been introduced to evaluate hand hygiene. However, current two-stage methods, which separate segmentation and Action Quality Assessment, suffer from unstable evaluation due to the reliance on the models for each respective task. Additionally, single-stage methods perform poorly in long-duration videos due to extensive background noise interference. Moreover, existing hand hygiene datasets have low annotation density and limited samples. To alleviate these limitations in methodology and data, this paper proposes a Single-stage Dual-task Joint Learning (SDJL) framework and a hand hygiene evaluation dataset (HHA1009). First, this model utilizes the vision transformer with the fixed token as the backbone, significantly improving computational efficiency and prediction accuracy by jointly learning Action Segmentation and quality assessment tasks. Additionally, the HHA1009 dataset captures video sequences in various scenarios and the individual action in each video is scored by multiple raters. The experimental results on the HHA1009 show the superior performance of the proposed SDJL, ablation studies further confirm the effectiveness of each module.