Interpretable Convolutional Neural Network for Violence Recognition
摘要
Video surveillance violence recognition is a sensitive task, and the lack of interpretability in ‘black-box’ models presents a major concern. As a result, efforts to enhance prediction interpretability in deep learning models for violence recognition are paramount. Therefore, this paper presents a case-based reasoning approach using prototype feature maps to improve the interpretability of a 3D neural network for violence recognition in surveillance footage. The deep learning model incorporates interpretability directly within its architecture, providing natural explanations for its decisions. The model’s predictions are based on the inputs and their nearest prototypes. The prototype feature maps also highlight active regions that the model uses to make its prediction, allowing for more fine-grained analysis. The results demonstrate that this approach enhances model interpretability while simultaneously achieving state-of-the-art performance on three publicly available surveillance violence recognition benchmark datasets: RWF-2000, SCFD, and ViolentFlows.