Task-adaptive debiasing with SCM for sentiment analysis
摘要
Sentiment analysis models often mirror social regularities found in data, which can lead to unfair or unreliable predictions. We address this challenge with a task adaptive debiasing method guided by the Stereotype Content Model, using the Warmth and Competence dimensions as signals. For each training example we compute SCM scores from a validated lexicon and set instance specific adversarial weights, so that examples with stronger stereotypical cues receive more debiasing pressure while neutral cases are largely preserved. We evaluate this approach on diverse text domains including product reviews, short movie phrases, and social media posts. Alongside standard accuracy, we monitor distributional parity using common fairness diagnostics such as the Demographic Parity gap and Kolmogorov–Smirnov