ALF-YOLO: a modified YOLOv8n algorithm for precise emotion detection via facial expressions
摘要
Detecting facial expressions plays a vital role in medical rehabilitation and human–computer interaction. This work proposes the attention-enhanced lightweight fusion (ALF-YOLO) framework, a modified version of YOLOv8n, to address challenges such as limited precision in emotion categorizing and inadequate real-time performance in traditional methods. The investigation introduces the adaptive multi-branch convolution (AMConv) module as an alternate for particular typical convolutional modules in the original model, thus improving its ability to recognize multi-scale facial information and complicated details. The large separable kernel attention (LSK-Attention) module is integrated into the backbone to emphasize critical spatial features, enhancing the model’s sensitivity to facial information. The interlayer feature injection (IFI) structure is also demonstrated, dramatically decreasing information loss and optimizing the incorporation of multi-scale face features, therefore, improving the model’s emotion recognition capability. Findings from experiments on the AffectNet dataset demonstrate that the proposed ALF-YOLO model achieves an optimal balance between real-time performance and precision compared with the original YOLOv8n. The model improves the mean average precision at 50% (mAP50) by 3.7%, achieves a detection speed of 282 FPS, and has a parameter size of roughly 2.81 M. The precision and generalization ability of the ALF-YOLO model has been demonstrated on the challenging BCCD and Mar20 datasets. The ALF-YOLO framework provides a successful approach for precision emotion detection by face expression.