A Multilingual Text Classification Debiasing Framework Based on Feature-Weighted Adversarial Prompt Tuning
摘要
Compared with the research on the fairness of monolingual text classification models, the research on the fairness of multilingual text classification models has broader universality. However, existing research has disregarded the importance of the features in text representations that lead to bias. These features include identity terms and language styles associated with sensitive attributes, which contribute significantly to classification tasks. In order to address this issue, we introduce a multilingual text classification debiasing framework based on feature-weighted adversarial prompt tuning. The adversarial training in this framework focuses on the biased features that contribute most to classification, enabling deeper identification and mitigation of bias. Additionally, prompt tuning is employed to conserve computational and memory resources. Experimental results show that, compared to existing methods, this framework achieves better accuracy-fairness trade-off.