Harmful Videos Classification for Safer Online Social Network
摘要
In this study, we introduce our self-collected dataset, HarmfulVideoVN2023 including 1,589 videos that are several seconds long, and investigate the effectiveness of image, audio, and language features for classifying assumedly harmful videos, particularly for children according to our defined labels. We apply 3D CNN models to video frames, transfer learning models for audio spectrograms and transformers for textual features. Our experimental findings indicate that training models using image frames decoded from videos demands substantial computational resources, our best-recorded result is 80.77% accuracy. In contrast, employing pre-trained models for images extracted from audio spectrograms and textual features derived from video yields accurate results and significantly reduces training time as well as computational cost (83.34% and 87.18% accuracy for audio spectrograms and texts respectively).