Violence Detection in Indoor Domestic Environment Using Multimodal Information
摘要
The recognition of violent activities has emerged as a prominent topic within the realm of computer vision research. This issue is of paramount importance within indoor domestic environments, given the heightened vulnerability of individuals residing in such conditions to potential attacks. Consequently, there exists a substantial imperative for the development of intelligent surveillance systems capable of autonomously monitoring individuals and identifying instances of violent behavior. While numerous techniques grounded in handcrafted and deep learning features have been proposed to address this challenge, it is noteworthy that many of these methodologies primarily focus on video data and often disregard the potential insights offered by audio information. In light of this, this paper presents a fusion model for violence detection, which integrates audio and video features to provide a comprehensive and nuanced assessment of potentially violent incidents. Our proposed approach leverages inflated 3D networks and pretrained models from the Visual Geometry Group to extract video and audio features, respectively. Subsequently, these extracted feature vectors are combined utilizing a joint cross-attention fusion model. The results of our proposed model are promising, with an achieved accuracy rate of 73.3%, coupled with a processing speed of 25 frames per second. This research contribution is instrumental in advancing violence detection technology, particularly within indoor domestic environments, where the safety and well-being of individuals are paramount concerns.