Developing a vision-based approach for identifying crowd panic in video surveillance systems is a complex task due to the struggle to gather enough real-world event recordings for training. The use of synthetic data can mitigate this issue, but the domain gap between synthetic and real-world samples needs to be managed to achieve precise results. We present a method to train these systems effectively by combining synthetic and real data to differentiate between normal and panic states. Our method learns domain-invariant spatio-temporal visual cues of the scenes along with supplementary descriptive attributes of crowd directions for the panic state classification. Experimental results show its potential with respect to alternative state-of-the-art methodologies and how it can effectively leverage synthetic data to train this kind of systems with high accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning Domain-Invariant Spatio-Temporal Visual Cues for Video-Based Crowd Panic Detection

  • Javier Calle,
  • Luis Unzueta,
  • Peter Leskovsky,
  • Jorge García

摘要

Developing a vision-based approach for identifying crowd panic in video surveillance systems is a complex task due to the struggle to gather enough real-world event recordings for training. The use of synthetic data can mitigate this issue, but the domain gap between synthetic and real-world samples needs to be managed to achieve precise results. We present a method to train these systems effectively by combining synthetic and real data to differentiate between normal and panic states. Our method learns domain-invariant spatio-temporal visual cues of the scenes along with supplementary descriptive attributes of crowd directions for the panic state classification. Experimental results show its potential with respect to alternative state-of-the-art methodologies and how it can effectively leverage synthetic data to train this kind of systems with high accuracy.