Compared with the research on the fairness of monolingual text classification models, the research on the fairness of multilingual text classification models has broader universality. However, existing research has disregarded the importance of the features in text representations that lead to bias. These features include identity terms and language styles associated with sensitive attributes, which contribute significantly to classification tasks. In order to address this issue, we introduce a multilingual text classification debiasing framework based on feature-weighted adversarial prompt tuning. The adversarial training in this framework focuses on the biased features that contribute most to classification, enabling deeper identification and mitigation of bias. Additionally, prompt tuning is employed to conserve computational and memory resources. Experimental results show that, compared to existing methods, this framework achieves better accuracy-fairness trade-off.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multilingual Text Classification Debiasing Framework Based on Feature-Weighted Adversarial Prompt Tuning

  • Yongmei Zhou,
  • Rongxi Zhou,
  • Dong Zhou,
  • Aimin Yang

摘要

Compared with the research on the fairness of monolingual text classification models, the research on the fairness of multilingual text classification models has broader universality. However, existing research has disregarded the importance of the features in text representations that lead to bias. These features include identity terms and language styles associated with sensitive attributes, which contribute significantly to classification tasks. In order to address this issue, we introduce a multilingual text classification debiasing framework based on feature-weighted adversarial prompt tuning. The adversarial training in this framework focuses on the biased features that contribute most to classification, enabling deeper identification and mitigation of bias. Additionally, prompt tuning is employed to conserve computational and memory resources. Experimental results show that, compared to existing methods, this framework achieves better accuracy-fairness trade-off.