An edge-optimized dual-teacher knowledge distillation framework with wavelet alignment for real-time still-image activity recognition
摘要
Still-image human activity recognition (HAR) has improved in recent years, but many existing models remain too heavy for real-time use on embedded or edge devices. This paper introduces an efficient dual-teacher knowledge distillation (KD) framework with a Vision Transformer (ViT) Teacher and an EfficientNet-B3 Teacher-Assistant to train a lightweight Student model. We apply wavelet-guided feature decomposition to capture multi-scale structures and use a linear projection to map the Teacher, Assistant, and Student features into a shared latent space for alignment. In addition, we propose a simple adaptive loss-weighting method called Wavelet-Residual-Guided Exponentiated Gradient(WRG–EG), which adjusts the KD loss components during training to improve stability and generalization. Experiments on the Stanford-40 and Willow Action datasets show that the proposed method reaches mean Average Precision (mAP) scores of 96.99 and 96.24, respectively. It also achieves a