Extracting distinctive features from crowds has been a key focus in crowd activity recognition. While the use of image and monocular video data has been prominent in crowd activity recognition, the utilization of stereo vision presents significant advantages, particularly in the acquisition of three-dimensional information. However, extracting and representing crowd features from binocular videos present challenges. This study presents a novel approach to crowd activity recognition using deep learning, which captures temporal changes in individual behavior while preserving core variations. The method also accounts for the mutual influence of individual behaviors through information diffusion and accumulation. By integrating long short-term memory networks and graph convolutional neural networks, the proposed method extracts spatiotemporal features of crowds, leveraging binocular videos to capture individual behavior and associations. Experimental results demonstrate that the model achieves superior recognition performance by utilizing crowd features derived from binocular videos.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Video Crowd Activity Recognition Based on Stereo Vision

  • Gang Zhang,
  • Cong Wang,
  • Yiwei Hu

摘要

Extracting distinctive features from crowds has been a key focus in crowd activity recognition. While the use of image and monocular video data has been prominent in crowd activity recognition, the utilization of stereo vision presents significant advantages, particularly in the acquisition of three-dimensional information. However, extracting and representing crowd features from binocular videos present challenges. This study presents a novel approach to crowd activity recognition using deep learning, which captures temporal changes in individual behavior while preserving core variations. The method also accounts for the mutual influence of individual behaviors through information diffusion and accumulation. By integrating long short-term memory networks and graph convolutional neural networks, the proposed method extracts spatiotemporal features of crowds, leveraging binocular videos to capture individual behavior and associations. Experimental results demonstrate that the model achieves superior recognition performance by utilizing crowd features derived from binocular videos.