<p>Intelligent monitoring systems often struggle with accurate human detection and action recognition in complex environments such as classrooms. To address this, we propose an improved human behavior recognition framework based on a modified YOLOv11 architecture. A key contribution of this study is the creation of the Student Classroom Behavior dataset (SCB-dataset3), a novel benchmark comprising 5686 images and 45,578 annotations across six behavior classes (hand-raising, reading, writing, phone interaction, head-bowing, desk-leaning) and twelve educational stages from preschool to university. Our model integrates the CBAM attention module and a dual classification head to enhance feature representation and enable simultaneous location and action classification. Optimization via quantization and pruning further boosts deployment efficiency. Experimental evaluations show that our model achieves a mean average precision (<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="44163_2025_492_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="71" /> </InlineMediaObject> <EquationSource Format="TEX">\({\text{mAP}}@0.{5}\)</EquationSource> </InlineEquation>) of 0.805 for human detection and 0.722 for action recognition, with an inference speed of 58.2 FPS, outperforming the YOLOv8n and baseline YOLOv11n models by 1.4 and 3.6% in <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="44163_2025_492_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="71" /> </InlineMediaObject> <EquationSource Format="TEX">\({\text{mAP}}@0.{5}\)</EquationSource> </InlineEquation>, respectively, while halving the parameter count compared to the dual YOLOv11 models. These results demonstrate superior performance in both accuracy and efficiency, validating the effectiveness of SCB-dataset3 and the proposed architecture for robust classroom behavior analysis.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A human location and action recognition method based on improved Yolov11 model

  • Shunyi Chen,
  • Yongkang Liu,
  • Hanqing Zhang,
  • Yi Cai

摘要

Intelligent monitoring systems often struggle with accurate human detection and action recognition in complex environments such as classrooms. To address this, we propose an improved human behavior recognition framework based on a modified YOLOv11 architecture. A key contribution of this study is the creation of the Student Classroom Behavior dataset (SCB-dataset3), a novel benchmark comprising 5686 images and 45,578 annotations across six behavior classes (hand-raising, reading, writing, phone interaction, head-bowing, desk-leaning) and twelve educational stages from preschool to university. Our model integrates the CBAM attention module and a dual classification head to enhance feature representation and enable simultaneous location and action classification. Optimization via quantization and pruning further boosts deployment efficiency. Experimental evaluations show that our model achieves a mean average precision ( \({\text{mAP}}@0.{5}\) ) of 0.805 for human detection and 0.722 for action recognition, with an inference speed of 58.2 FPS, outperforming the YOLOv8n and baseline YOLOv11n models by 1.4 and 3.6% in \({\text{mAP}}@0.{5}\) , respectively, while halving the parameter count compared to the dual YOLOv11 models. These results demonstrate superior performance in both accuracy and efficiency, validating the effectiveness of SCB-dataset3 and the proposed architecture for robust classroom behavior analysis.