A benchmark dataset for evaluating gender sensitivity in Korean political discourse with large language models
摘要
Large language models are increasingly applied to political discourse, but their ability to detect culturally grounded gender sensitivity remains underexplored. We introduce KOGENT, a benchmark dataset of 1,222 transcripts from the Korean National Assembly, annotated for gender sensitivity across 6,024 utterances. Each utterance is labeled as high or low in gender sensitivity, based on contextual indicators of bias, discrimination, or inclusion, and tagged for the target group. KOGENT spans Korean legislative sessions from 1948 to 2024. Annotation reliability was ensured through dual coding and adjudication, yielding high intercoder agreement. When tasked with labeling utterances by gender sensitivity, GPT-4.1 achieved F1-scores of 87.5% (zero-shot) and 91.2% (18-shot), while GPT-4o reached 90.4% and 91.1%, respectively. While incorporating in-domain examples enhanced model performance, limitations in distinguishing between criticisms and reinforcements of inequality, culturally specific terminology, and extended contexts were observed for both models. Our results demonstrate KOGENT’s utility as a robust benchmark for analyzing gender sensitivity in Korean political speech and evaluating multilingual LLMs’ sociocultural alignment.