<p>Language shapes our thoughts and perceptions, influencing concepts like gender roles. Biased language has personal and societal consequences, enforcing exclusion, affecting labor force participation, reinforcing stereotypes, and deepening social inequalities. Detecting biased language is essential yet challenging; manual review is time-consuming and subjective. Existing machine learning (ML) and natural language processing solutions for gender bias face limitations like insufficient definitions, taxonomies, and annotated data, often relying on simplistic statistical analyses that fail to capture contextual sentence-level biases. To address these gaps, we propose Genderly, a hybrid and modular data-centric system for detecting gender biases in English across linguistic levels. Genderly identifies gender-specific terms, analyzes idiomatic expressions, phrases, and metaphors perpetuating biases, and captures implicit biases within nuanced language elements like sarcasm and figurative speech. Our contributions include designing Genderly, a proof-of-concept tool with a user feedback loop for ongoing improvement. We also curated new datasets as community resources for bias detection and mitigation. Our experiments used diverse preprocessing, feature engineering, and hyperparameter optimization methods for traditional ML models and large language models (LLMs) as gender bias detectors, comparing these results to evaluate model performance. Error analysis further explored top-performing model strengths and limitations. Findings reveal that optimized ML models detect surface-level biases effectively, while LLMs excel in identifying biases embedded in semantics, with benevolent sexism being particularly challenging. By helping users recognize biases, Genderly promotes fairness and inclusive language. Through continuous improvements, Genderly aims for equitable language use and a more inclusive society.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Genderly: a data-centric gender bias detection system

  • Wael Khreich,
  • Jad Doughman

摘要

Language shapes our thoughts and perceptions, influencing concepts like gender roles. Biased language has personal and societal consequences, enforcing exclusion, affecting labor force participation, reinforcing stereotypes, and deepening social inequalities. Detecting biased language is essential yet challenging; manual review is time-consuming and subjective. Existing machine learning (ML) and natural language processing solutions for gender bias face limitations like insufficient definitions, taxonomies, and annotated data, often relying on simplistic statistical analyses that fail to capture contextual sentence-level biases. To address these gaps, we propose Genderly, a hybrid and modular data-centric system for detecting gender biases in English across linguistic levels. Genderly identifies gender-specific terms, analyzes idiomatic expressions, phrases, and metaphors perpetuating biases, and captures implicit biases within nuanced language elements like sarcasm and figurative speech. Our contributions include designing Genderly, a proof-of-concept tool with a user feedback loop for ongoing improvement. We also curated new datasets as community resources for bias detection and mitigation. Our experiments used diverse preprocessing, feature engineering, and hyperparameter optimization methods for traditional ML models and large language models (LLMs) as gender bias detectors, comparing these results to evaluate model performance. Error analysis further explored top-performing model strengths and limitations. Findings reveal that optimized ML models detect surface-level biases effectively, while LLMs excel in identifying biases embedded in semantics, with benevolent sexism being particularly challenging. By helping users recognize biases, Genderly promotes fairness and inclusive language. Through continuous improvements, Genderly aims for equitable language use and a more inclusive society.