QLGWYB: design of an efficient model for analyzing crowd behavior through Quad LSTM and Quad GRU fusion enhanced by Q-learning and YOLO
摘要
In the field of computer vision, analyzing Crowd Behavior poses significant challenges due to the complexity of human dynamics in dense populations. Traditional methods often struggle to accurately detect and classify individual actions within crowds, leading to limited effectiveness in behavior analysis. This study introduces a novel approach by integrating You Only Look Once Version 8 (YOLOv8), Q-learning, Quad Long Short-Term Memory (Quad LSTM), and Quad Gated Recurrent Units (Quad GRU) to address these challenges. Our methodology begins with YOLOV8 for initial human detection, followed by Q-learning to enhance You Only Look Once detection accuracy. This combination improves detection efficiency and facilitates the detailed classification of actions. The integration of Quad LSTM and Quad GRU enables our model to extract and analyze complex motion features from different directions, allowing for a deeper understanding of both individual and collective behaviors. The model was evaluated using the Text Retrieval Conference Video Retrieval Evaluation (TRECVID) and Crowd 11 datasets, achieving significant improvements over existing methods: a 5.9% increase in precision, 4.5% in accuracy, 4.9% in recall, 8.3% in processing speed, 9.5% in Area Under the Curve (AUC), and 8.5% in behavioral specificity. These advancements not only offer a more precise tool for analyzing Crowd Behavior but also have potential applications in public safety, urban planning, and surveillance. Proposed approach sets a new standard in interpreting human dynamics, marking a significant contribution to future research in computer vision and Crowd Behavior analysis.