Multi-sensor Data Fusion for Accurate Prediction of Suspicious Activities in Video
摘要
Video analytics for surveillance enhances safety of the public across a range of public spaces, such as shopping centers, railroad stations, and daycares. Conventional closed-circuit television systems have proven to be inefficient and prone to errors as they rely heavily on human operators. The proposed intelligent video surveillance system utilizes deep learning algorithms to promptly notify the authorities of any suspicious activity. The system has been trained to identify a wide range of human actions and potential hazards. It achieves this by utilizing datasets like UCF-Crime, HMDB51, and UCF101. The architecture combines LSTM (Long Short Term Memory) with CNN (Convolution Neural Network). The CNN component effectively identifies suspicious activity by extracting important features from video frames using the Inception V3 model. The LSTM examines these characteristics to establish connections over time. The system demonstrated exceptional performance on UCF-Crime, with an F1 score of 99.1%, precision of 99.8%, recall of 98.5%, and an impressive accuracy of 99.6%, as observed during the evaluation across datasets. HMDB51 achieves an impressive F1 score of 99.1%, with precision at 98.7%, recall at 97.9%, and accuracy at 99.4%. On the other hand, UCF101 achieves a commendable F1 score of 98.6%, with accuracy at 99.2%, precision at 98.7%, and recall at 97.9%. The detection durations for UCF-Crime, HMDB51, and UCF101 are 0.25 s, 0.30 s, and 0.28 s, respectively. The proposed model streamlines the procedure of improving safety by utilizing intelligent video analytics.