<p>Bias in Artificial Intelligence (AI) systems is widely recognized as a critical challenge across technical, educational, and societal domains, yet its conceptual treatment frequently remains anchored at the level of observable model outputs. This conflation of symptoms with causes obscures the layered, interdependent nature of how bias arises and propagates across the AI cycle, and limits the effectiveness of both mitigation strategies and instruction. This paper proposes a framework of ten sources of harm in AI, organized across three analytically distinct but interdependent levels: dataset and source bias, model and training bias, and interaction and surface bias. The third level constitutes the primary contribution of this work, introducing three sources of harm that arise specifically at the human-AI interaction layer and are absent from concise AI bias frameworks: presentation bias, the distortion of user perception through interface ranking, linguistic hedging, and visual salience; framing bias, the systematic sensitivity of model outputs to surface-level variations in prompt formulation independent of semantic content; and interaction bias, the emergent amplification of bias through iterative user engagement and its feedback into future training data. The taxonomy extension is developed following the method of Nickerson et al. [<CitationRef CitationID="CR22">22</CitationRef>] and evaluated against their ending conditions. A formal treatment of the framework as a series of functional transformations across the AI cycle enables precise localization of each bias type to a distinct injection point. Because the ten categories are not interchangeable descriptions of the same underlying problem, the choice of remediation depends critically on identifying where in the pipeline a distortion originates—applying a remedy suited to one category to a bias originating in another displaces rather than neutralizes the harm. For engineering practice, the framework supports structured bias audits in which localization precedes remediation. For educational practice, it provides a principled basis for curricula and educational materials in which individual bias mechanisms are taught in isolation before being encountered in combination.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bias in AI: a framework for understanding the sources of bias in artificial intelligence

  • Wolfgang Robinig,
  • Johannes P. Wallner

摘要

Bias in Artificial Intelligence (AI) systems is widely recognized as a critical challenge across technical, educational, and societal domains, yet its conceptual treatment frequently remains anchored at the level of observable model outputs. This conflation of symptoms with causes obscures the layered, interdependent nature of how bias arises and propagates across the AI cycle, and limits the effectiveness of both mitigation strategies and instruction. This paper proposes a framework of ten sources of harm in AI, organized across three analytically distinct but interdependent levels: dataset and source bias, model and training bias, and interaction and surface bias. The third level constitutes the primary contribution of this work, introducing three sources of harm that arise specifically at the human-AI interaction layer and are absent from concise AI bias frameworks: presentation bias, the distortion of user perception through interface ranking, linguistic hedging, and visual salience; framing bias, the systematic sensitivity of model outputs to surface-level variations in prompt formulation independent of semantic content; and interaction bias, the emergent amplification of bias through iterative user engagement and its feedback into future training data. The taxonomy extension is developed following the method of Nickerson et al. [22] and evaluated against their ending conditions. A formal treatment of the framework as a series of functional transformations across the AI cycle enables precise localization of each bias type to a distinct injection point. Because the ten categories are not interchangeable descriptions of the same underlying problem, the choice of remediation depends critically on identifying where in the pipeline a distortion originates—applying a remedy suited to one category to a bias originating in another displaces rather than neutralizes the harm. For engineering practice, the framework supports structured bias audits in which localization precedes remediation. For educational practice, it provides a principled basis for curricula and educational materials in which individual bias mechanisms are taught in isolation before being encountered in combination.