Identifying predictive factors of cyberbullying perpetration via deep learning: a two-stage training approach to class imbalance with weighted loss
摘要
Although machine learning has been widely used to investigate problematic behaviors such as cyberbullying, class imbalance in behavioral datasets remains a persistent challenge. This study, based on the General Aggression Model and the Social-Ecological Framework, aims to improve the identification of cyberbullying perpetrators through a deep learning framework that combines a two-stage training strategy with a class-weighted loss function.
MethodsA total of 660 Chinese university students (including 54 self-reported perpetrators) were recruited from schools.
ResultsThe deep learning model achieved a recall of 0.92, significantly outperforming conventional methods such as LightGBM and Random Forest which showed low recall for the minority class. SHAP analysis revealed that screen time, negative emotions, school connectedness, online disinhibition, exposure to violent media and deviant peer affiliation were the most influential predictors. Interaction analyses revealed that school connectedness may buffer the impact of multiple risk factors.
ConclusionsThe framework effectively addresses extreme class imbalance for cyberbullying detection, offering both strong performance and interpretable insights for prevention.