<p>Representational harms influence people’s views, attitudes, and perceptions of particular social groups. In journalism, the task of captioning images is particularly susceptible to such harms, as journalists must balance accuracy with sensitivity in their descriptions. Poorly written captions can perpetuate stereotypes or marginalize certain groups. This study explores how Large Language Models can be utilized to minimize representational harms in image captions. The primary objective is to deploy prompting techniques to generate captions that minimize potential harms, aligned with established scientific frameworks. To achieve this, we have developed three prompting strategies: Baseline Response generates unfiltered captions as a benchmark; Reactive Mitigation employs a feedback mechanism that revises initial captions using examples of harmful content; and Proactive Guidance integrates definitions and examples of harmful content prior to generating captions to guide the model toward more contextually aware outputs. We compared the captions generated by these strategies to those created by professional journalists, finding that AI-generated captions often minimize representational harms more effectively. Our evaluation indicated that Reactive Mitigation emerged as the most effective strategy for reducing potential harms. These findings highlight the importance of a socio-technical approach, informed by the Social Construction of Technology (SCOT) theory, to combines the efficiency of automated tools with the depth of journalistic expertise, ultimately fostering more accurate representation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Navigating representation: utilizing prompt engineering to minimize representational harms in journalist’s image captions

  • Habiba Sarhan,
  • Morteza Shahrezaye,
  • Simon Hegelich

摘要

Representational harms influence people’s views, attitudes, and perceptions of particular social groups. In journalism, the task of captioning images is particularly susceptible to such harms, as journalists must balance accuracy with sensitivity in their descriptions. Poorly written captions can perpetuate stereotypes or marginalize certain groups. This study explores how Large Language Models can be utilized to minimize representational harms in image captions. The primary objective is to deploy prompting techniques to generate captions that minimize potential harms, aligned with established scientific frameworks. To achieve this, we have developed three prompting strategies: Baseline Response generates unfiltered captions as a benchmark; Reactive Mitigation employs a feedback mechanism that revises initial captions using examples of harmful content; and Proactive Guidance integrates definitions and examples of harmful content prior to generating captions to guide the model toward more contextually aware outputs. We compared the captions generated by these strategies to those created by professional journalists, finding that AI-generated captions often minimize representational harms more effectively. Our evaluation indicated that Reactive Mitigation emerged as the most effective strategy for reducing potential harms. These findings highlight the importance of a socio-technical approach, informed by the Social Construction of Technology (SCOT) theory, to combines the efficiency of automated tools with the depth of journalistic expertise, ultimately fostering more accurate representation.