Exposing shortcut sensitivity in English–Norwegian hate speech detection via neutral stability analysis
摘要
The spread of hate speech has become a serious issue in online communities, negatively influencing user interaction and platform integrity. Although major social media companies have devoted substantial effort to developing automated detection systems, reliably classifying hateful content continues to be difficult. One key reason is shortcut learning, where models rely on frequent or sensitive words instead of understanding sentence meaning and context. Many hate speech detection systems achieve high accuracy while relying on surface-level lexical cues rather than robust semantic understanding. This shortcut-based behavior can lead to unstable predictions, particularly for neutral texts containing identity-related terms. In this paper, we analyze the prediction stability of transformer-based hate speech models when classifying neutral content. We construct a bilingual dataset of English and Norwegian social media posts annotated for binary hate speech detection and evaluate several Norwegian-specific and multilingual transformer models under identical training conditions. To mitigate shortcut learning, we introduce a neutral only stability regularization objective that encourages consistent predictions under stochastic perturbations. In post-training, we assess residual instability by applying controlled masking-based input variations through random token masking and measuring prediction variability. Our results show that Norwegian-specific models, especially Nor-BERT